כתבה
arXiv cs.LG ·
Logit Refiner: שיפור דגימות תמונה אוטומטיות באמצעות דגימת תלות פנימית
Logit Refiner: Improving Visual Autoregressive Models via Intra-Scale Dependency Modeling
אנו מציגים את Logit Refiner, שיפור דגימות תמונה אוטומטיות באמצעות דגימת תלות פנימית. השיפור נעשה על ידי דגימת תקני תלות פנימית, שמשפרת את תפוקת הדגימה. השיפור נבדק על מודל LangGraph.
תקציר מקורי באנגליתarXiv:2609.11804v1 Announce Type: cross Abstract: Visual Autoregressive Models (VAR) generate images through next-scale prediction, producing all tokens within each scale in parallel. We show that this parallel decoding constitutes a mean-field-style approximation that discards spatial dependencies among same-scale tokens, causing locally incoherent samples regardless of backbone capacity -- a limitation of the decoding rule. Addressing this limitation, we introduce the Logit Refiner, a lightweight autoregressive module that restores intra-scale dependencies by sequentially sampling tokens conditioned on frozen backbone features. Adding only ~10% parameters and less than 5% of the base model's training compute, it plugs into any pretrained VAR checkpoint without retraining. Controlled abla
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית