כתבה
arXiv cs.AI ·
סגירת מעגל: מתכונים מעשיים לאימון מודלי שפה מחוברים
Closing the Loop: Practical Training Recipes for Looped Language Models
חוקרים פיתחו מתכונים מעשיים לאימון מודלי שפה מחוברים. המחקר מראה כי ניתן לאמן מודלים כאלה בעלות נמוכה יותר ולשפר את ביצועיהם. הם השוו את המודל LoopLM למודל צפוף ומצאו שהוא משפר את הביצועים במגוון משימות.
תקציר מקורי באנגליתarXiv:2610.00673v1 Announce Type: cross Abstract: Looped language models increase effective depth by repeatedly applying a shared block of layers, but existing large-scale recipes require multi-stage training over trillions of tokens, while the benefits of recurrence remain difficult to separate from differences in data and training. In this work, we establish practical training recipes for looped language models, with three main results. (1) We develop a compute-efficient from-scratch pipeline that reduces the training budget from 7.7T tokens in Ouro to 310B tokens while retaining strong reasoning performance. Pretraining followed by high-quality mid-training, together with learning-rate warmup and stronger exit-gate regularization, enables stable recurrent training without prior multi-st
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית