כתבה
arXiv cs.AI ·
DIAL-OPD: Learning More from Fewer Tokens in On-Policy Distillation
תקציר מקורי באנגליתarXiv:2610.11659v1 Announce Type: cross Abstract: On-policy distillation (OPD) supervises student-generated trajectories with token-level teacher signals. Its sampled-token variant avoids the cost of full-vocabulary probabilities. Yet we find that training on fewer tokens can outperform full-token OPD, challenging the intuition that more supervision improves learning. This motivates selecting tokens by learning value. Existing disagreement-based criteria ignore probability scale: tokens assigned negligible probability by both models, termed low-low tokens, can receive large log-ratio rewards and hinder learning. We propose DIAL-OPD, a token-selection method that bridges log-probability and probability spaces by weighting reward magnitude with the logarithmic mean of teacher and student pro
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית