יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

התפשטות כמות-פרובביליסטית: רוטאציה של התפשטות על-מדיני

Distillation as Probability Transport: Routed On-Policy Distillation
אופצייה של התפשטות על-מדינית (OPD) עוברת ידע של המורה למדל של התלמיד, אך יעדים נבחרים יעילים מפחיתים את ההפצה של המורה לסקלר של זכות על פריט ספציפי. כזכות זו מעידה האם פריט יצטרך לקבל או לאבד סבירות, אך פונדקאות המתאימה אינה ניתונה. אנו חולקים OPD כהפעלת-כוח של הפעלת-כוח של פרובביליסטיקה ומציעים RouteOPD (Routed On-Policy Distillation), שמפרקת אי-ספיקה מקומית של המורה-התלמיד למקורות-עודף של התלמיד וליעדים-חסרים של המורה, ומקשרת אותם לזוגות-הפעלה-כוח. RouteOPD משפרת סטטיסטיקה-לוג-אודס כלפי יעדים-משותפים שניתנים להשגה מפוטנציאל-מורה-מוגבל, ומתאימה את תקציב ההפעלה-כוח לצפיפות דרישה-מורה. תפיסה זו מכוונת את העדכונים ליעדים-מועדפים של המורה ומשליטה את גודלם בתוך פעלת-כוח-הפעלה-כוח אחת.
תקציר מקורי באנגליתarXiv:2609.08337v1 Announce Type: new Abstract: On-policy distillation (OPD) transfers teacher knowledge on student-generated trajectories, but efficient sampled objectives reduce the teacher distribution to scalar credit on individual tokens. Such credit indicates whether a token should gain or lose probability, yet leaves the corresponding redistribution unspecified. We recast OPD as teacher-guided probability transport and propose RouteOPD (Routed On-Policy Distillation), which decomposes local teacher--student disagreement into student-excess sources and teacher-deficit destinations and couples them into explicit transport pairs. RouteOPD optimizes pairwise log-odds toward jointly realizable targets obtained from a bounded teacher potential, while adapting the transport budget to the c
קרא במקור המקורי