כתבה
arXiv cs.LG ·
EasyPPO: Stabilizing the Critic Is Key
תקציר מקורי באנגליתarXiv:2609.36802v1 Announce Type: new Abstract: A key strength of Proximal Policy Optimization (PPO) is its learned critic, which uses historical trajectories collected during reinforcement learning to estimate expected returns and reduce policy-gradient variance. However, we find that the critic is also a major source of instability in reinforcement learning for large language models (LLMs). We identify two critic failure modes that destabilize PPO. First, filtering truncated rollouts from both actor and critic shifts the policy objective to reward conditioned on completion, allowing truncation to increase even as conditional reward improves. Second, heterogeneous return noise can cause high-variance prompts to dominate critic updates in finite batches. We introduce EasyPPO to address the
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית