כתבה
arXiv cs.AI ·
RELACE: retrospective likelihood-based action credit estimation for long-horizon language agents
תקציר מקורי באנגליתarXiv:2610.07349v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) avoids a separate critic by estimating advantages from rollout groups. For multi-turn agents, however, trajectory-level supervision provides coarse, noisy credit: terminal rewards do not locate errors and can penalize useful actions alongside mistakes. Group-in-Group Policy Optimization (GiGPO) and subsequent methods refine supervision through state-conditioned comparisons, but their credit estimates remain sensitive to downstream decisions and outcomes. We introduce RELACE, Retrospective Likelihood-based Action, a critic-free framework that integrates retrospective action assessment with state-conditioned advantage estimation. RELACE evaluates executed actions through teacher-forced likelihood scor
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית