כתבה
arXiv cs.AI ·
On the Off-Policy Teacher in On-Policy Distillation
פתרון לבעיית האסימטריה בהטמעה על-מדורגת: SCOUT, פרקטיקה של עדכון משותף של המורה
תקציר מקורי באנגליתarXiv:2609.38360v1 Announce Type: cross Abstract: On-policy distillation (OPD) has recently emerged as a promising post-training paradigm in which the student learns from trajectories generated by its own policy under dense teacher supervision. However, OPD introduces a fundamental asymmetry: although the sampled trajectories are on-policy for the student, they are off-policy for the teacher. The teacher is typically optimized to continue from prefixes generated by its own policy, but during OPD it must instead supervise prefixes generated by the student. Empirically, we find that its continuation performance degrades as these prefixes grow longer. To address this issue, we propose Student-COnditioned Updates of the Teacher (SCOUT), a co-training framework that adapts the teacher to studen
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית