יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

השפעות קבועות של אימון על-טבעי מעבר לבלבול

Lasting Effects of Abstract Pretraining Beyond Perplexity
אימון על-טבעי קודם לאימון טבעי משפר את יכולות המודל הלשוני מעבר לבלבול. נמצא שאימון קצר במשימה של עיבוד סטק על-טבעי, המחייב יכולות תיאורטיות, משפר את יכולות המודל לטיפול בשאלות רב-פסיפס. נמצא גם שאימון זה קורה רק אם המשימה העל-טבעית ניתנת קודם לאימון הטבעי, ולא במהלכו.
תקציר מקורי באנגליתarXiv:2609.38764v1 Announce Type: new Abstract: Language models are typically pretrained from random initialization. Recent work challenges this convention, showing that a brief warm-up on abstract, algorithmically generated data can provide a better starting point for subsequent learning of natural language. In this paper, we show that in small language models, such a warm-up improves specific capabilities that are not reflected in language-modeling perplexity. Our warm-up uses an abstract stack-manipulation task that requires compositional and state-tracking capabilities. Allocating as little as 1% of pretraining tokens to this data improves multi-hop question answering by up to 3.9 F1 points on MUSIQUE, with additional gains on HOTPOTQA and 2WIKIMULTIHOPQA despite comparable language-mo
קרא במקור המקורי