כתבה
arXiv cs.CL ·
Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillation
תקציר מקורי באנגליתarXiv:2609.39687v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) trains mathematical reasoning models using a privileged teacher that sees a reference solution and supervises student-sampled prefixes. Standard OPSD uses one fixed parameter setting at every state, but nearby settings may offer additional supervision. We find that local parameter perturbations reveal complementary reference-aligned corrections under the same reference context. Different experts supply these corrections at different reference positions. Their pool covers more such positions than the unperturbed privileged teacher. We introduce Neighborhood OPSD (N-OPSD) to turn these corrections into supervision at student-visited states. Offline, greedy selection builds a compact pool of frozen experts by r
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית