יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

Privy to the Foil: Recasting Value Estimation with a Self-Privileged Critic for RLVR

תקציר מקורי באנגליתarXiv:2609.37825v1 Announce Type: new Abstract: Assigning credit to intermediate steps remains a central challenge in training Large Language Models (LLMs) on multi-step reasoning tasks with sparse terminal rewards, and actor-critic methods such as PPO address this by learning value functions to construct token-level advantages. Their effectiveness, however, hinges on reliable value estimation, a difficult task requiring the critic to both assess progress toward a correct solution and anticipate an evolving policy's future behavior; errors in either can compromise credit assignment and destabilize online training. In this paper, we revisit the standard state-only formulation of value estimation and propose $\pi$PPO, a self-privileged actor-critic framework. By reusing verified same-prompt
קרא במקור המקורי