יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

מרכוז ערוצי פלט לאימון LLM יציב

Output Embedding Centering for Stable LLM Pretraining
חוקרים מציעים שיטה חדשה לייצוב אימון מודלי שפה גדולים. השיטה, הנקראת מרכוז ערוצי פלט, מטפלת בבעיית התפרקות הלוגיטים בסוף האימון. השיטה הוכחה כיעילה בהשוואה לשיטות אחרות.
תקציר מקורי באנגליתarXiv:2601.02031v3 Announce Type: replace-cross Abstract: Pretraining of large language models is not only expensive but also prone to certain training instabilities. A specific instability that often occurs at the end of training is output logit divergence. The most widely used mitigation strategies, z-loss and logit soft-capping, merely address the symptoms rather than the underlying cause of the problem. In this paper, we analyze the instability from the perspective of the output embeddings' geometry and identify anisotropic embeddings as its source. Based on this, we propose output embedding centering (OEC) as a new mitigation strategy, and demonstrate that it suppresses output logit divergence. OEC can be implemented in two different ways: as a deterministic operation called $\mu$-cen
קרא במקור המקורי