כתבה
arXiv cs.AI ·
E$^2$-OPSD: Taming Entropy Overshoot in On-Policy Self-Distillation
תקציר מקורי באנגליתarXiv:2610.05048v2 Announce Type: replace-cross Abstract: On-policy self-distillation (OPSD) provides dense token-level supervision without a second model: one network acts as teacher with the reference solution and as student with only the problem. We identify a specific failure mode of this recipe. During training, student token entropy rises past the teacher's and remains elevated, a pattern we call entropy overshoot. We trace it to both sides of distillation. The reference-conditioned teacher is confident along its answer-directed reasoning path, but this confidence transfers poorly to student-generated prefixes, making its supervision overly tied to answer-specific cues rather than reusable reasoning patterns; meanwhile, the forward KL used by OPSD continually diffuses the student's p
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית