כתבה
arXiv cs.LG ·
GraphOPD: גרפים מוגברים להטמעה של פוליציה-על-מדרגה לאג'נטים של LLM
GraphOPD: Graph-Augmented On-Policy Distillation for LLM Agents
GraphOPD משפר את הטמעה העל-מדרגה של פוליציה לאג'נטים של LLM על ידי שימוש בגרפים מוגברים של סטרוקטורלי אוגמנטציה. השיטה קוראת את הצעדים שאפשרו את הצעדים הבאים מתוך ההסברה של האקראוט, ומארגנת אותם לגרף של תלות. היא מחשבת את הקרדיט הסטרוקטורלי של כל צעד על ידי תפוצה סטציונרית רנדומית על הגרף, ומשלבת אותו עם התאריך הדיברגנסיה למספור של כל רולאוט. GraphOPD מציגה יכולת תחרותית כל פעם מחדש, ומשפרת על הבסיס החזק ביותר בעד 5.8%.
תקציר מקורי באנגליתarXiv:2610.08959v1 Announce Type: new Abstract: On-policy distillation post-trains large language model agents by supplying dense, step-level guidance from a teacher policy when the reinforcement-learning reward is sparse and arrives only once per trajectory. Existing instantiations allocate this guidance by the size of the teacher-student divergence at each step, on the single-turn intuition that a large disagreement marks a mistake worth correcting. Once decisions chain over many turns, that rule misfires, since an early drift enters every later context both policies condition on, leaving the teacher consistent with the drifted trajectory instead of flagging its cause, while interchangeable steps register large but outcome-irrelevant divergences. We demonstrate this on an agentic benchma
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית