כתבה
arXiv cs.AI ·
Solving Without Stopping: On-Policy Distillation at Small Scale
תקציר מקורי באנגליתarXiv:2609.37326v1 Announce Type: new Abstract: On-policy distillation, where a student learns from a stronger teacher's feedback on its own outputs, is a common way to pass reasoning to smaller models. We analyze what it transfers at small scale, distilling Qwen3-8B into Qwen3 4B, 1.7B and 0.6B students, in thinking mode (reason at length, then end the reasoning and answer) and, for comparison, in non-thinking mode (no separate reasoning phase). Long reasoning needs two abilities, solving a problem and knowing when it is solved, and we find that distillation transfers the first, but in thinking mode not the second. Solving improves at every size, up to two ceilings, which we measure comprehensively across both modes and all student sizes: a student's single attempt never exceeds what it c
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית