כתבה
arXiv cs.CL ·
Diagnosing On-Policy Self-Distillation for Reasoning Language Models
תקציר מקורי באנגליתarXiv:2609.39118v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) has attracted growing interest as a promising approach to improve the reasoning ability of language models. Without external rewards nor a separate stronger teacher, the self-teacher with privileged information could provide dense signals on student's trajectories. However, its behavior in language reasoning remains unclear, with reported outcomes ranging from modest gains to behavioral collapse. In this work, we diagnose OPSD for mathematical reasoning across models spanning 0.6B--8B parameters. We conduct controlled experiments and token-level analyses to fully delve into OPSD. We point out that teacher's signal is shaped by reasoning-mode alignment and the complete teacher prefix, rather than by privilege
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית