כתבה
arXiv cs.CL ·
לימוד לפעול עם התקדמות משימה
Learning to Act with Task Progress: Distilling Small Agents from Compact Teacher Supervision
שיטת Task-Progress Distillation מאפשרת אימון סוכנים קטנים לביצוע משימות חוזרות. השיטה משתמשת בהדרכה קומפקטית ומאפשרת לסוכנים ללמוד מהתקדמות המשימה. הניסויים הראו תוצאות טובות עם 404 הדגמות, ואף עלו על גישות אחרות במקרים מסוימים.
תקציר מקורי באנגליתarXiv:2610.10332v1 Announce Type: new Abstract: Learning from large-model demonstrations offers a way to train small agents that can complete recurring tasks without calling a large model at every step. A central design choice is what to retain from teacher trajectories that contain reasoning, actions, and information about task progress. We introduce Task-Progress Distillation (TPD), an offline approach that pairs each demonstrated action with a short label describing the current task stage. The student learns these compact targets and selects actions by jointly scoring admissible stage--action pairs, which a deterministic harness executes in the environment. On ALFWorld, a 1.7B student trained with 404 demonstrations achieves 72.4\% mean unseen task success with either TPD or action-only
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית