כתבה
arXiv cs.LG ·
אימות לפני ניסיון: שער תלמיד-מורה ברמת הפניה
Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation
TGOPD משפרת את האימון המוקדם על ידי בדיקת אמינות המורה. היא משתמשת בניסויים קטנים כדי לבדוק את האמינות של המורה לפני העברת הנתונים. TGOPD משפרת את התוצאות ב-6 סביבות שונות.
תקציר מקורי באנגליתarXiv:2609.02998v1 Announce Type: new Abstract: On-policy distillation (OPD) accelerates post-training by providing dense token-level supervision from a frozen teacher on the student's own rollouts. Vanilla OPD applies this supervision uniformly across prompts, without checking whether the teacher is reliable for each prompt. Because reverse KL is mode-seeking, a confidently wrong teacher can induce a strong yet misleading update. Distributional proxies, such as entropy or teacher-student likelihood agreement, measure uncertainty or agreement but do not directly verify outcome correctness. We introduce Teacher-Gated On-Policy Distillation (TGOPD), built on the principle that teacher reliability should be verified at the prompt level before dense supervision is admitted. TGOPD estimates rel
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית