כתבה
arXiv cs.CL ·
להבנת תהליך האימון של טוקן פאוז: נקודת מבט של תחזוקה
Towards Understanding Pause Token Fine-Tuning Dynamics: A Mode Retention Perspective
אימון טוקני פאוז משפר את ההסתברות של LLM על ידי שינוי תהליך האימון. נמצאו תוצאות של עד 6 נקודות במתמטיקה ו-2.5 נקודות בקוד.
תקציר מקורי באנגליתarXiv:2609.04489v1 Announce Type: new Abstract: Pause-token methods improve LLM reasoning by inserting special tokens into sequences. Prior work explains these gains through computational expressivity. However, there is relatively little investigation into the training dynamics of pause tokens. We explore how pause tokens reshape the training dynamics of fine-tuning. Two controlled pilots expose distinct asymmetries. On a synthetic continual-learning task, masked pauses overwrite a previously-learned distribution roughly 4x less at matched final adaptation (H1, mode retention); on a synthetic math-reasoning probe, the boundary-adjacent token comes to encode substantially more downstream-step information (H2, non-myopic compression). We formalize a training rule consistent with both - Maske
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית