יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

STEPS: Selective On-Policy Self-Distillation for Reasoning

תקציר מקורי באנגליתarXiv:2605.10194v2 Announce Type: replace Abstract: On-policy self-distillation (OPSD) uses a model as its own teacher under privileged context, providing token-level supervision on the model's own reasoning trajectories. However, persistent all-token guidance can degrade training in our math RL setting. Such guidance may unnecessarily constrain non-critical tokens and reinforce biases induced by privileged information unavailable at inference. Motivated by these risks, we propose STEPS, a selective OPSD framework that controls the location, direction, and duration of distillation. STEPS identifies critical spans in student rollouts, applies forward KL to key spans or reverse KL to error spans, and gradually transitions to pure GRPO after a short distillation phase. Across four models from
קרא במקור המקורי