כתבה
arXiv cs.LG ·
הבנת דיסטילציה של פוליסי-אוף vs. און-פוליסי: סיפור של מטרות אימון שונות
Understanding Off- vs On-Policy Distillation: A Tale of Distinct Training Objectives
במאמר זה, חוקרים חוקרים את יעילות והגמישות של דיסטילציה של פוליסי-אוף. הם חוקרים את השימוש בפוליסי-אוף כדי לשפר את יעילות האימון של רשתות עצביות.
תקציר מקורי באנגליתarXiv:2609.38666v1 Announce Type: new Abstract: On-policy distillation (OPD) learns from teacher feedback on student-generated responses and has shown promise in reducing forgetting relative to supervised fine-tuning (SFT). However, its benefits and fragility remain incompletely understood. We study sequential distillation from multiple teachers, where the student minimizes its average divergence from the teachers. Forward Kullback--Leibler (KL) divergence yields a weighted arithmetic mixture, while reverse KL yields a normalized weighted geometric aggregate. We develop algorithms that learn these targets under off-policy and on-policy feedback, respectively, establishing logarithmic regret bounds in the tabular setting and extending the analysis to function approximation. By analyzing the
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית