כתבה
arXiv cs.AI ·
התאמת הנפיצה, לא החישוב
Match the Distribution, Not the Compute: Post-Training Multi-Token Prediction Heads
חוקרים מציגים שיטה לשיפור ביצועי מודלי שפה על ידי התאמת ראשי ניבוי רב-טוקנים בשלב האימון. השיטה מאפשרת להגיע לביצועים דומים לאלו של מודלים שאומנו במשותף, אך עם פחות נתונים. החוקרים בדקו את השיטה על מודל Qwen3-8B והראו שהיא משיגה שיפורים משמעותיים בביצועים.
תקציר מקורי באנגליתarXiv:2610.00888v1 Announce Type: cross Abstract: Multi-token prediction (MTP) improves the throughput of autoregressive generation by enabling the language model to draft multiple next tokens per forward pass, while a verification step over draft tokens ensures that token distribution of the backbone is preserved. Every open MTP-family release (MiMo-7B, DeepSeek-V3, Qwen3) trains its heads jointly with the backbone over the full pretraining run of tens of trillions of tokens, thus setting the drafter quality at pretraining time. We ask whether a lightweight post-training pass on target-generated chain-of-thought is enough to reach the same expected throughput speedup on a frozen reasoning model, and study how a serving-time system built on such a checkpoint can be optimized. We present th
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית