כתבה
arXiv cs.CL ·
פתרון ללא הפסקה: דיסטילציה על-מדינה בקנה מידה קטן
Solving Without Stopping: On-Policy Distillation at Small Scale
במאמר זה, המחברים מבצעים הדגמה של דיסטילציה על-מדינה בקנה מידה קטן, כדי לבדוק מה נעבר לדגמים קטנים. הם משתמשים במודל Qwen3-8B כמורה חזקה, ומדגימים את תהליך הדיסטילציה לדגמים קטנים יותר. התוצאות המוצגות במאמר עוסקות באופן שבו דגמים קטנים מתנהגים תחת דיסטילציה על-מדינה, ובאופן שבו ניתן להבחין בין תכונות שונות של דגמים קטנים.
תקציר מקורי באנגליתarXiv:2609.37326v1 Announce Type: cross Abstract: On-policy distillation, where a student learns from a stronger teacher's feedback on its own outputs, is a common way to pass reasoning to smaller models. We analyze what it transfers at small scale, distilling Qwen3-8B into Qwen3 4B, 1.7B and 0.6B students, in thinking mode (reason at length, then end the reasoning and answer) and, for comparison, in non-thinking mode (no separate reasoning phase). Long reasoning needs two abilities, solving a problem and knowing when it is solved, and we find that distillation transfers the first, but in thinking mode not the second. Solving improves at every size, up to two ceilings, which we measure comprehensively across both modes and all student sizes: a student's single attempt never exceeds what it
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית