כתבה
arXiv cs.LG ·
שיכול אונ-פוליסי עם רשתות תלויות גרף
Graph-Conditioned On-Policy Agent Distillation from Off-the-Shelf Teachers
GC-OPD הוא שיטה חדשה לאימון סוכנים קומפקטים עם מורים מוכנים מראש. השיטה משתמשת ברשתות תלויות גרף כדי לשפר את הביצועים של הסוכנים. GC-OPD הראתה שיפורים משמעותיים במשימות שונות, כולל ScienceWorld ו-ALFWorld.
תקציר מקורי באנגליתarXiv:2609.37522v2 Announce Type: replace Abstract: On-policy distillation (OPD) trains compact language agents with teacher feedback on student-generated trajectories. In multi-turn tasks, compounding errors can move students beyond the teacher's effective supervision. We introduce Graph-Conditioned On-Policy Agent Distillation (GC-OPD), which enriches an off-the-shelf teacher's scoring context with execution evidence. A graph indexes repeated teacher executions by shared states while preserving complete successful and failed histories. After each student episode, GC-OPD retrieves current-state references or historical alternatives and combines them with student hindsight to score the original thought-action tokens. Using the same original teachers, GC-OPD improves mean success over vanil
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית