כתבה
arXiv cs.LG ·
TISD: On-Policy Self-Distillation with Trajectory Intervention
תקציר מקורי באנגליתarXiv:2609.30878v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) provides dense teacher targets, but evaluates them only along student-sampled rollouts. When the privileged teacher favors an alternative action at a visited prefix, OPSD can provide a target for the branch decision but cannot supervise the successor contexts induced by that action unless the student samples it. This creates a training-time data-collection bottleneck and suggests a different role for teacher-student disagreement: proposing a trajectory branch rather than identifying a sufficient local repair. Our diagnostic framework using controlled token interventions reveals that a teacher-preferred token at peak disagreement can improve student continuation success, while its local corrective value is
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית