כתבה
arXiv cs.LG ·
מה עוזר לחזור בעצמו בדגמי שפה?
What Makes Recurrence Effective in Looped Language Models?
דגמי שפה שמחזירים את עצמם מצליחים לשפר את חישובי ההשערה ללא הוספת תכונות.
תקציר מקורי באנגליתarXiv:2609.36636v1 Announce Type: new Abstract: Looped language models (LoopLMs) increase computational depth through parameter sharing, offering a path to scale inference computation without adding parameters. However, it remains unclear when additional recurrence is beneficial and how architectural choices affect its effectiveness. Through controlled experiments, we systematically examine (1) when recurrence helps, (2) where it should be applied, and (3) how its conditioning affects performance. Our evaluation covers inference budgets below, within, and beyond the training horizon under knowledge and reasoning tasks. (1) We find that recurrence can improve reasoning beyond the training horizon while degrading knowledge performance, but harder reasoning instances do not consistently benef
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית