כתבה
arXiv cs.CL ·
למד לעצמך איפה לבקר: עיבוד עצמי של תשומת לב על-מדיקטיבי לסיבוכיות
Teach Yourself Where to Look: On-Policy Attention Self-Distillation for Reasoning
עיבוד עצמי על-מדיקטיבי מאמן מודלי סיבוכיות על נתיביהם העצמיים. המאמר מציג טכניקה חדשה של עיבוד עצמי של תשומת לב, שמספקת סיגנל הדרכה נוסף שמשפר את דיוק, יציבות וכושר חישוב של העיבוד העצמי.
תקציר מקורי באנגליתarXiv:2609.33200v2 Announce Type: replace-cross Abstract: On-policy self-distillation trains reasoning models on their own trajectories using dense token distribution guidance from a privileged teacher with access to a verified solution. This supervision transfers what the teacher predicts without directly transferring where it attends within the preceding context. We introduce On-Policy Attention Self-Distillation (OPASD), which complements token-level supervision with solution-conditioned attention distillation. Because the privileged teacher can attend to verified solution tokens unavailable to the student, OPASD projects teacher attention onto student-visible positions and renormalizes the resulting distribution before alignment. Across three model sizes and four competition-level math
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית