יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

שיפור עיבוד תוצאות עם מורה

Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability
חוקרים בודקים את השפעת הרחבת ההתאמה על תהליך הלמידה. הם מוצאים כי הרחבת ההתאמה אינה משפרת את הדיוק. המחקר מציע גישה חדשה לניתוח הנתונים.
תקציר מקורי באנגליתarXiv:2610.08448v1 Announce Type: new Abstract: On-Policy Distillation (OPD) trains a student on its own generations using teacher feedback. With different tokenizers, comparing teacher and student predictions requires alignment at both sequence and vocabulary levels. In this paper, we examine whether expanding this alignment coverage improves learning. Across three heterogeneous teacher--student pairs on mathematical reasoning and code generation, strict 1:1 groups already cover most student-generated tokens despite substantial vocabulary mismatch. On responses sampled from the students before distillation, the shared vocabulary retains nearly all teacher and student probability mass at strictly aligned positions on average. Restricting reverse KL to a student-selected top-16 subset of th
קרא במקור המקורי