יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

דינמיקת עדינות טוקני פסקת עדינה

Towards Understanding Pause Token Fine-Tuning Dynamics: A Mode Retention Perspective
חוקרים פיתחו שיטה חדשה לשיפור יכולות ההיגיון של מודלים Qwen ו-Llama. השיטה, הנקראת Masked Boundary Pause, משתמשת בטוקני פסקה כדי לשפר את יכולות ההיגיון. החוקרים מצאו שהשיטה משפרת את היכולות ההיגיוניות של המודלים, תוך שמירה על יכולות השפה הכלליות.
תקציר מקורי באנגליתarXiv:2609.04489v1 Announce Type: cross Abstract: Pause-token methods improve LLM reasoning by inserting special tokens into sequences. Prior work explains these gains through computational expressivity. However, there is relatively little investigation into the training dynamics of pause tokens. We explore how pause tokens reshape the training dynamics of fine-tuning. Two controlled pilots expose distinct asymmetries. On a synthetic continual-learning task, masked pauses overwrite a previously-learned distribution roughly 4x less at matched final adaptation (H1, mode retention); on a synthetic math-reasoning probe, the boundary-adjacent token comes to encode substantially more downstream-step information (H2, non-myopic compression). We formalize a training rule consistent with both - Mas
קרא במקור המקורי