כתבה
arXiv cs.LG ·
Graph-Conditioned On-Policy Agent Distillation from Off-the-Shelf Teachers
תקציר מקורי באנגליתarXiv:2609.37522v1 Announce Type: new Abstract: On-policy distillation (OPD) trains compact language agents with teacher feedback on student-generated trajectories. In multi-turn tasks, compounding errors can move students beyond the teacher's effective supervision. We introduce Graph-Conditioned On-Policy Agent Distillation (GC-OPD), which enriches an off-the-shelf teacher's scoring context with execution evidence. A graph indexes repeated teacher executions by shared states while preserving complete successful and failed histories. After each student episode, GC-OPD retrieves current-state references or historical alternatives and combines them with student hindsight to score the original thought-action tokens. Using the same original teachers, GC-OPD improves mean success over vanilla O
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית