כתבה
arXiv cs.AI ·
מתכון מסולסל למומחים מחוברים
Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts
LOOM הוא מתכון חדש למודלים מסולסלים של מומחים מחוברים. הוא מאפשר הגדלת עומק המודל מבלי להגדיל את מספר הפרמטרים. הניסויים הראו שיפור בביצועים עם 5-9 לולאות.
תקציר מקורי באנגליתarXiv:2610.01153v1 Announce Type: cross Abstract: Looped Transformers introduce recurrent depth as a new scaling axis for LLMs: by repeatedly applying shared Transformer blocks, they increase effective depth without increasing parameter count. However, the benefits of looping remain unclear for large MoE LLMs under FLOPs-matched comparisons. The main reason is that the gains from additional iterations diminish quickly and can even turn into degradation, so the extra FLOPs spent on looping yield little substantial improvement. Consequently, prior work typically settles on two loops. We identify two main obstacles to scaling looped MoE. First, looping inherits and amplifies the curse of depth: hidden-state variance grows with each iteration as residual updates accumulate, which destabilizes
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית