כתבה
arXiv cs.CL ·
CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning
תקציר מקורי באנגליתarXiv:2609.36820v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) is widely used to train reasoning language models, where it computes advantages by centering and normalizing rewards across rollouts of the same prompt. For multiple rewards, GRPO sums the reward components and normalizes the total reward by its within-group standard deviation. The corresponding variance equals the sum of all pairwise reward covariances. For a fixed centered reward, larger aggregate covariance produces smaller advantages, and vice versa, allowing update magnitudes to adapt to reward dependence. However, correlated rewards with large scales can dominate this normalization and suppress signals from smaller-scale rewards. We propose Correlation-Normalized GRPO (CorrGRPO), which normali
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית