כתבה
arXiv cs.LG ·
Towards Better Training Signal: Advantage Clipped Policy Optimization
תקציר מקורי באנגליתarXiv:2609.36816v1 Announce Type: new Abstract: Reinforcement learning (RL) has become a cornerstone for improving the reasoning capabilities of large language models (LLMs), but the need for on-policy data substantially limits training efficiency. Reusing off-policy data through importance sampling (IS) can improve efficiency but introduce considerable instability. Hence, algorithms such as PPO and GRPO widely adopt IS-ratio clipping to stabilize training. However, training stability and gradient estimate are mainly determined by the product of IS ratio and advantage. To further stabilize training, we propose ACPO, which clips the product of the IS ratio and the advantage, leading to more stable gradient estimates. We also establish a connection between ACPO and gradient clipping in polic
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית