יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

DIAL-OPD: למידה מטוקנים מועטים

DIAL-OPD: Learning More from Fewer Tokens in On-Policy Distillation
DIAL-OPD הוא שיטה לבחירת טוקנים בתהליך אילוף מודלים. היא משפרת את ביצועי האילוף על ידי בחירה מושכלת של טוקנים. DIAL-OPD מוכיחה עליונות על שיטות אחרות בתחום.
תקציר מקורי באנגליתarXiv:2610.11659v1 Announce Type: new Abstract: On-policy distillation (OPD) supervises student-generated trajectories with token-level teacher signals. Its sampled-token variant avoids the cost of full-vocabulary probabilities. Yet we find that training on fewer tokens can outperform full-token OPD, challenging the intuition that more supervision improves learning. This motivates selecting tokens by learning value. Existing disagreement-based criteria ignore probability scale: tokens assigned negligible probability by both models, termed low-low tokens, can receive large log-ratio rewards and hinder learning. We propose DIAL-OPD, a token-selection method that bridges log-probability and probability spaces by weighting reward magnitude with the logarithmic mean of teacher and student proba
קרא במקור המקורי