כתבה
arXiv cs.CL ·
Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning
תקציר מקורי באנגליתarXiv:2609.36178v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment of trajectory-level advantages to all policy tokens fails to distinguish consequential decisions from less relevant ones, obscuring which intermediate decisions contributed to success. We introduce ProVer, a framework that targets potentially pivotal decisions for fine-grained credit assignment in agentic reinforcement learning. Given a rollout group, an agentic judge contrasts successful and failed trajectories to propose a segment potentially responsible for their divergent outcomes. Rather than directly trusting the judge's assessment, ProVer verifies the proposed segment by estimating its adva
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית