יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

SAPD: Step-Aligned Privileged Distillation

תקציר מקורי באנגליתarXiv:2610.09665v1 Announce Type: new Abstract: On-policy post-training can improve large language models by learning from their own trajectories, but requires costly rollout generation. We ask whether fixed demonstrations can support competitive off-policy learning through better supervision. Our premise is that their usefulness depends not only on the training trajectories, but also on whether supervision provides informative preferences among continuations and connects this guidance to the reasoning decision being learned. We introduce Step-Aligned Privileged Distillation (SAPD), a rollout-free self-distillation method that turns demonstrations into step-aligned distributional supervision. Its key insight is to use the known progression of a reference solution to associate each reasonin
קרא במקור המקורי