כתבה
arXiv cs.AI ·
SR-OPSD: Self-Referenced On-Policy Self-Distillation
תקציר מקורי באנגליתarXiv:2608.09745v2 Announce Type: replace-cross Abstract: On-policy self-distillation (OPSD) converts feedback into dense token-level supervision on student-generated trajectories, complementing reinforcement learning with sparse outcome rewards. Its self-teacher, derived from the student's current or exponentially averaged parameters and conditioned on additional context, evolves alongside the student and its rollout context distribution. The benefit of modifying this moving target depends on how target--student probability mismatches translate into updates. We propose \emph{Self-Referenced On-Policy Self-Distillation (SR-OPSD)}, which constructs a normalized geometric target from the self-teacher and a frozen initial policy, then minimizes the forward R\'enyi divergence from this target
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית