כתבה
arXiv cs.AI ·
מגביר את המודלים האוטורגרסיביים הנורמליים: תיקון לוגיט
Logit Refiner: Improving Visual Autoregressive Models via Intra-Scale Dependency Modeling
מגבירים את המודלים האוטורגרסיביים הנורמליים: תיקון לוגיט. ניתן לשפר את האיכות של הדוגמאות שנוצרות על ידי המודלים האוטורגרסיביים הנורמליים על ידי תיקון לוגיט.
תקציר מקורי באנגליתarXiv:2609.11804v1 Announce Type: cross Abstract: Visual Autoregressive Models (VAR) generate images through next-scale prediction, producing all tokens within each scale in parallel. We show that this parallel decoding constitutes a mean-field-style approximation that discards spatial dependencies among same-scale tokens, causing locally incoherent samples regardless of backbone capacity -- a limitation of the decoding rule. Addressing this limitation, we introduce the Logit Refiner, a lightweight autoregressive module that restores intra-scale dependencies by sequentially sampling tokens conditioned on frozen backbone features. Adding only ~10% parameters and less than 5% of the base model's training compute, it plugs into any pretrained VAR checkpoint without retraining. Controlled abla
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית