כתבה
arXiv cs.LG ·
REVO: Rollout-Efficient Off-Policy Distillation via Variance-Guided Reuse
תקציר מקורי באנגליתarXiv:2609.37500v1 Announce Type: new Abstract: On-policy distillation (OPD) trains language models using dense token-level teacher supervision on student-generated trajectories. However, its reliance on frequently refreshed student rollouts often incurs substantial generation cost. We introduce REVO, an off-policy distillation framework that improves rollout efficiency by reusing each student rollout for multi-step learner updates. REVO addresses prefix-level and current-token policy mismatch through stabilized prefix weighting and one-step resampling from the current student, which enables repeated updates without regenerating full trajectories. To prioritize informative token positions within reused rollouts, REVO uses the variance of the student-teacher log-probability ratio to quantif
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית