יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

TIGPO: אופטימיזציה של מדיניות גרפית מקומית לסוכנים LLM

TIGPO: Temporal Instance-Graph Policy Optimization for Long-Horizon LLM Agents
TIGPO היא שיטה חדשה לאופטימיזציה של מדיניות גרפית מקומית עבור סוכנים LLM. היא מאפשרת שימוש במידע היסטורי כדי לשפר את הערכת היתרון היחסי. TIGPO נבדקה על ALFWorld ו-WebShop והראתה תוצאות טובות יותר משיטות קודמות.
תקציר מקורי באנגליתarXiv:2609.03383v1 Announce Type: new Abstract: Graph-based policy optimization improves credit assignment for long-horizon LLM agents by organizing rollout trajectories into state-transition graphs. However, existing methods construct graphs independently within each policy update, discarding transitions discovered by earlier policies and limiting advantage estimation to small, batch-local rollout groups. We propose \emph{Temporal Instance-Graph Policy Optimization} (TIGPO), which extends graph-based credit assignment across policy updates. TIGPO maintains a persistent transition graph for each task, allowing valid transitions discovered by different policy versions to jointly determine credit for current rollouts. To actively reconnect current exploration with historical experience, TIGP
קרא במקור המקורי