יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

הבחנה בין דיסטילציה על-פי-מדיניות לדיסטילציה על-פי-מדיניות: סיפור של מטרות אימון מסוגים שונים

Understanding Off- vs On-Policy Distillation: A Tale of Distinct Training Objectives
במאמר זה, חוקרים חוקרים את דיסטילציה על-פי-מדיניות ודיסטילציה על-פי-מדיניות, ומציעים אלגוריתמים ללמידה של יעדים אלה תחת פידבק על-פי-מדיניות ופידבק על-פי-מדיניות.
תקציר מקורי באנגליתarXiv:2609.38666v1 Announce Type: cross Abstract: On-policy distillation (OPD) learns from teacher feedback on student-generated responses and has shown promise in reducing forgetting relative to supervised fine-tuning (SFT). However, its benefits and fragility remain incompletely understood. We study sequential distillation from multiple teachers, where the student minimizes its average divergence from the teachers. Forward Kullback--Leibler (KL) divergence yields a weighted arithmetic mixture, while reverse KL yields a normalized weighted geometric aggregate. We develop algorithms that learn these targets under off-policy and on-policy feedback, respectively, establishing logarithmic regret bounds in the tabular setting and extending the analysis to function approximation. By analyzing t
קרא במקור המקורי