כתבה
arXiv cs.AI ·
איך טרנספורמרים לומדים לתכנן דרך תחזוקת מסרים רב-טוקן
How Transformers Learn to Plan via Multi-Token Prediction
במחקר זה נחקרה תכונת הלמידה של טרנספורמרים לתכנן דרך תחזוקת מסרים רב-טוקן. התוצאות הראו שטכניקה זו יעילה יותר מתכונת התחזוקת הטוקן הבא.
תקציר מקורי באנגליתarXiv:2604.11912v2 Announce Type: replace-cross Abstract: While next-token prediction (NTP) has been the standard objective for training language models, it often struggles to capture global structure in reasoning tasks. Multi-token prediction (MTP) has recently emerged as a promising alternative, yet its underlying mechanisms remain poorly understood. In this paper, we study how MTP facilitates reasoning, with a focus on planning. Empirically, we show that MTP consistently outperforms NTP on both synthetic graph path-finding tasks and more realistic reasoning benchmarks, such as Countdown and boolean satisfiability problems. Theoretically, we analyze a simplified two-layer Transformer on a star graph task. We prove that MTP induces a two-stage reverse reasoning process: the model first at
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית