כתבה
arXiv cs.AI ·
שיפור תהליך התכלסות במודלים
Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability
חוקרים בודקים את היעילות של שיטות חדשות לתכלסות מודלים. הם מצאו כי הרחבת הכיסוי של התכלסות אינה תמיד משפרת את התוצאות. החוקרים הציעו שיטה חדשה המתמקדת באמינות הפיקוח ולא בכיסוי מרבי.
תקציר מקורי באנגליתarXiv:2610.08448v1 Announce Type: cross Abstract: On-Policy Distillation (OPD) trains a student on its own generations using teacher feedback. With different tokenizers, comparing teacher and student predictions requires alignment at both sequence and vocabulary levels. In this paper, we examine whether expanding this alignment coverage improves learning. Across three heterogeneous teacher--student pairs on mathematical reasoning and code generation, strict 1:1 groups already cover most student-generated tokens despite substantial vocabulary mismatch. On responses sampled from the students before distillation, the shared vocabulary retains nearly all teacher and student probability mass at strictly aligned positions on average. Restricting reverse KL to a student-selected top-16 subset of
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית