כתבה
arXiv cs.LG ·
Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL
תקציר מקורי באנגליתarXiv:2610.00574v1 Announce Type: new Abstract: Multi-reward reinforcement learning trains large language models to satisfy multiple behavioral objectives simultaneously. Reward-wise normalization, as used in GDPO, preserves reward-specific relative information within rollout groups, but different objectives can still exhibit uneven learning progress. We study this behavior through advantage energy, the sum of a reward's squared advantages over a batch. Under idealized GDPO normalization, we show that this energy is proportional to active-group density: the fraction of rollout groups in which the reward provides nonzero relative advantages. This reveals a residual batch-level signal imbalance and provides a basis for calibrating reward contributions. Based on this relation, we propose Dens
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית