כתבה
arXiv cs.LG ·
CADENCE: Closing the Reasoning Gap via Coverage-Adaptive On-Policy Distillation
תקציר מקורי באנגליתarXiv:2607.16955v1 Announce Type: new Abstract: On-policy knowledge distillation transfers reasoning from large teachers to compact students, but existing approaches suffer three compounding failure modes: (i) cold-start collapse, where a fresh student assigns near-zero mass to teacher-preferred tokens; (ii) state-agnostic divergence scheduling, where time-only forward/reverse-KL interpolation ignores the student's coverage state; and (iii) binary reward sparsity, where pass/fail signals discard information from partially correct traces. We present CADENCE, a unified framework with a targeted fix for each. Its DRIFT mechanism schedules a per-token convex mixture of forward-KL and reverse-KL surrogate objectives on student-sampled trajectories (per-token surrogates, not sequence-level KL gr
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית