יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

התאמת הנפיצה, לא החישוב

Match the Distribution, Not the Compute: Post-Training Multi-Token Prediction Heads
חוקרים בדקו האם ניתן לשפר את מהירות הייצור של מודלים כמו Qwen3 ו-DeepSeek-V3 באמצעות התאמת ראשי ניבוי רב-טוקנים בשלב האימון. הם מצאו שניתן להשיג שיפורים משמעותיים במהירות הייצור תוך שימוש בפחות מ-1% מכמות הנתונים הנדרשת לאימון משותף.
תקציר מקורי באנגליתarXiv:2610.00888v1 Announce Type: new Abstract: Multi-token prediction (MTP) improves the throughput of autoregressive generation by enabling the language model to draft multiple next tokens per forward pass, while a verification step over draft tokens ensures that token distribution of the backbone is preserved. Every open MTP-family release (MiMo-7B, DeepSeek-V3, Qwen3) trains its heads jointly with the backbone over the full pretraining run of tens of trillions of tokens, thus setting the drafter quality at pretraining time. We ask whether a lightweight post-training pass on target-generated chain-of-thought is enough to reach the same expected throughput speedup on a frozen reasoning model, and study how a serving-time system built on such a checkpoint can be optimized. We present thre
קרא במקור המקורי