כתבה
arXiv cs.LG ·
RELACE: חישוב זכויות פעולה על יסוד של סבירות רטרוספקטיבית לשם חיזוי פעולות ארוכות-טווח
RELACE: retrospective likelihood-based action credit estimation for long-horizon language agents
מסגרת חינם-קריטית לחישוב זכויות פעולה באגנטים של שפה באורך-טווח. RELACE משלב חישוב סבירות רטרוספקטיבית עם הערכת יתרונות, ומציע זכויות פעולה דק-גרסה שמשלבות תגמול והערכת יתרונות. המסגרת נבחנה באגנט Qwen2.5-1.5B-Instruct ו-Qwen2.5-7B-Instruct, והראתה שיפורים ניכרים על פני GRPO, GiGPO ו-HCAPO.
תקציר מקורי באנגליתarXiv:2610.07349v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) avoids a separate critic by estimating advantages from rollout groups. For multi-turn agents, however, trajectory-level supervision provides coarse, noisy credit: terminal rewards do not locate errors and can penalize useful actions alongside mistakes. Group-in-Group Policy Optimization (GiGPO) and subsequent methods refine supervision through state-conditioned comparisons, but their credit estimates remain sensitive to downstream decisions and outcomes. We introduce RELACE, Retrospective Likelihood-based Action, a critic-free framework that integrates retrospective action assessment with state-conditioned advantage estimation. RELACE evaluates executed actions through teacher-forced likelihood scorin
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית