כתבה
arXiv cs.AI ·
TISD: עצמי-היסקה עם התערבות מסלול
TISD: On-Policy Self-Distillation with Trajectory Intervention
TISD הוא אלגוריתם חדש שמשפר את תהליך ההיסקה העצמית באמצעות התערבות מסלול. הוא מאפשר למורה להציע פעולות חלופיות ולבטל את הצורך בדגימה מחדש. TISD מראה שיפורים משמעותיים בתוצאות לעומת שיטות קודמות.
תקציר מקורי באנגליתarXiv:2609.30878v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) provides dense teacher targets, but evaluates them only along student-sampled rollouts. When the privileged teacher favors an alternative action at a visited prefix, OPSD can provide a target for the branch decision but cannot supervise the successor contexts induced by that action unless the student samples it. This creates a training-time data-collection bottleneck and suggests a different role for teacher-student disagreement: proposing a trajectory branch rather than identifying a sufficient local repair. Our diagnostic framework using controlled token interventions reveals that a teacher-preferred token at peak disagreement can improve student continuation success, while its local corrective value is li
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית