כתבה
arXiv cs.AI ·
Guide, Then Let Go: Gap-Adaptive Teacher Scheduling for Sparse-Reward Agentic RL
המאמר עוסק בפיתוח תוכנה של תוכנת RL שמסתגלת לקשיי הפעלה ומשפיעה על תהליך הלמידה. התוכנה נקראת Gap-Adaptive Teacher Scheduling (GATS) והיא משתמשת בשיטת on-policy distillation (OPD) כדי לספק הנחיות לתלמיד. התוכנה נבחנה במספר תפריטי תצוגה והיא הציגה תוצאות טובות יותר מאשר שיטות אחרות.
תקציר מקורי באנגליתarXiv:2609.37898v1 Announce Type: new Abstract: Reinforcement learning for long-horizon agents typically relies on sparse outcome-based rewards. This leads to a severe cold-start problem, as early-stage policies often fail to solve sampled tasks, leaving little useful reward signal for learning. To mitigate this problem, we use on-policy distillation (OPD) to provide token-level guidance on the student's own rollouts. We find that the benefit of this guidance depends on the performance gap between the teacher and the student. When the teacher substantially outperforms the student, distillation helps guide the student through the early training stage where outcome rewards provide little learning signal. As the gap narrows and eventually reverses, however, continued distillation becomes less
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית