כתבה
arXiv cs.LG ·
Interpolated Policy Distillation: A Controllable Continuum Between Off-Policy and On-Policy Distillation
תקציר מקורי באנגליתarXiv:2609.37170v1 Announce Type: new Abstract: Off-policy and on-policy distillation have traditionally been formulated as separate paradigms, each favoring a different property of distillation trajectories. Teacher-generated (off-policy) traces are typically high-quality but lie far from the student's distribution, whereas student-generated (on-policy) rollouts are more learnable but often contain erroneous reasoning. We view these paradigms as the endpoints of a policy continuum and posit that a more effective rollout policy may lie in between. We introduce \textbf{Interpolated Policy Distillation (IPD)}, which defines the next-token distribution at every decoding step as an explicit linear interpolation between the student and teacher distributions. The interpolation operates at the di
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית