כתבה
arXiv cs.LG ·
טרנספורמרים לומדים לתכנ� אתרים באמצעות ניבוי רב-טוקנים
How Transformers Learn to Plan via Multi-Token Prediction
חוקרים בדקו כיצד טרנספורמרים לומדים לתכנן אתרים באמצעות ניבוי רב-טוקנים. הם מצאו שניבוי רב-טוקנים משפר את היכולת לתכנן אתרים בהשוואה לניבוי טוקן בודד. המחקר מראה כיצד ניבוי רב-טוקנים מאפשר לטרנספורמרים ללמוד ולתכנן אתרים בצורה יותר יעילה.
תקציר מקורי באנגליתarXiv:2604.11912v2 Announce Type: replace Abstract: While next-token prediction (NTP) has been the standard objective for training language models, it often struggles to capture global structure in reasoning tasks. Multi-token prediction (MTP) has recently emerged as a promising alternative, yet its underlying mechanisms remain poorly understood. In this paper, we study how MTP facilitates reasoning, with a focus on planning. Empirically, we show that MTP consistently outperforms NTP on both synthetic graph path-finding tasks and more realistic reasoning benchmarks, such as Countdown and boolean satisfiability problems. Theoretically, we analyze a simplified two-layer Transformer on a star graph task. We prove that MTP induces a two-stage reverse reasoning process: the model first attends
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית