כתבה
arXiv cs.LG ·
Why On-Policy Distillation Sometimes Fails: Vanishing Learning Signals
תקציר מקורי באנגליתarXiv:2610.11247v1 Announce Type: new Abstract: On-policy distillation (OPD) enables effective capability transfer between language models, yet the mechanisms underlying its failures are not fully understood. Across code generation and mathematical reasoning, OPD with larger-scale teachers exhibits early loss plateaus, with an average final loss reduction of 25.1% after 200 updates, compared with 96.2% for self-RL teachers, obtained by further reinforcement learning (RL) training of the initial student. To understand this difference, we analyze OPD as an idealized continuous-time dynamical system in the small-learning-rate limit. Our training-log diagnostics associate these plateaus with an early decline in a gradient-based learning-signal proxy while substantial loss remains; these measur
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית