כתבה
arXiv cs.AI ·
התערבות מורה סטוכסטית להיסטוריה של מדיניות
Stochastic Teacher Intervention for Agentic On-Policy Distillation
חוקרים פיתחו שיטה חדשה להעברת ידע ממודל מורה חזק למודל תלמיד חלש יותר. השיטה, STI-OPD, משתמשת בהתערבות מורה סטוכסטית כדי לשפר את תהליך האימון. החוקרים בדקו את השיטה במספר מערכות ומצאו שהיא משפרת את ביצועי המודל התלמיד.
תקציר מקורי באנגליתarXiv:2610.10878v1 Announce Type: cross Abstract: On-policy distillation (OPD) efficiently transfers capabilities from a stronger teacher to a student language model through dense token-level supervision on student-generated rollouts and has shown promise on complex tasks such as mathematical reasoning. However, in multi-turn agentic tasks, student decisions shape subsequent observations, causing early errors to accumulate across turns. The resulting trajectories can drift away from the teacher's rollout distribution, making the teacher's token-level supervision less reliable or even counterproductive for OPD training. To address this issue, we introduce STI-OPD, a stochastic teacher intervention framework for multi-turn agentic OPD. During multi-turn interaction, STI-OPD uses teacher inte
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית