יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

מהפכה מבין דיסוננס לאורכסטרציה: התערבות מורה בהתפשטות על-פוליציה

From Dissonance to Orchestration: Teacher Intervention in On-Policy Distillation
התערבות מורה משפרת התפשטות על-פוליציה, אך עומק ומיקום אופטימליים משתנים. המאמר מציג את MAESTRO, שמשתמש בסקור של מחלוקת מדיניות כדי לשפר את ההתערבות של המורה. MAESTRO הציג תוצאות טובות יותר מאשר שיטות אחרות בשמונה מבחני סיבוכיות מתמטיים.
תקציר מקורי באנגליתarXiv:2609.37510v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student on its own reasoning trajectories using feedback from a stronger teacher. Teacher interventions can improve these trajectories, but also change the distribution on which the student learns. Our controlled studies show that rollout quality alone is an incomplete criterion for allocating teacher guidance. Deeper intervention yields diminishing gains in rollout accuracy while increasing off-policy load. In a training probe with a restricted rollout horizon, peak student accuracy and performance retention favor different intervention strengths. The preferred intervention depth and placement also vary across benchmarks. These findings motivate MAESTRO, which uses local policy disagreement to jointly ad
קרא במקור המקורי