כתבה
arXiv cs.CL ·
התערבות מורה-סטוכסטית להעברת יכולות באופן-מדירה
Stochastic Teacher Intervention for Agentic On-Policy Distillation
התערבות מורה-סטוכסטית להעברת יכולות באופן-מדירה. פיתוח של OPD למשימות אגנטיות רב-פניות.
תקציר מקורי באנגליתarXiv:2610.10878v1 Announce Type: new Abstract: On-policy distillation (OPD) efficiently transfers capabilities from a stronger teacher to a student language model through dense token-level supervision on student-generated rollouts and has shown promise on complex tasks such as mathematical reasoning. However, in multi-turn agentic tasks, student decisions shape subsequent observations, causing early errors to accumulate across turns. The resulting trajectories can drift away from the teacher's rollout distribution, making the teacher's token-level supervision less reliable or even counterproductive for OPD training. To address this issue, we introduce STI-OPD, a stochastic teacher intervention framework for multi-turn agentic OPD. During multi-turn interaction, STI-OPD uses teacher interv
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית