כתבה
arXiv cs.LG ·
פיענוח טרנספורמרים מחוברים טוב יותר (כמעט) בחינם
Decoding Looped Transformers Better for (Almost) Free
LoopCD הוא כלי פיענוח קונטרסטיבי שמשפר את יכולת הפיענוח של טרנספורמרים מחוברים. הוא עובד עם מודלים כמו LLaMA ומסגרות כמו LangChain. LoopCD מאפשר שיפור באיכות הפיענוח תוך הפחתה משמעותית של עומס החישוב.
תקציר מקורי באנגליתarXiv:2610.02185v1 Announce Type: new Abstract: Looped Transformers achieve parameter efficiency by repeatedly executing a shared block across recurrent loops. Each loop yields an intermediate representation decodable for the same next token, yet standard decoding discards earlier states. Because earlier loops embody less computation, recurrence inherently supplies aligned weak-and-strong prediction pairs without auxiliary models or external training. We introduce LoopCD, a training-free contrastive decoding framework that guides token selection by contrasting the final prediction with an earlier recurrent pass, operating either in logit space with one extra output pass (LoopCD-Logits) or in hidden-state space with zero output overhead (LoopCD-Hidden). Across four looped Transformer famili
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית