כתבה
arXiv cs.LG ·
GRPODropout: Less is More for Online Reinforcement Learning Rollouts
תקציר מקורי באנגליתarXiv:2610.11854v1 Announce Type: new Abstract: Reinforcement learning (RL) methods such as GRPO substantially improve large language model reasoning but often suffer from policy entropy collapse: the loss of sampling diversity weakens exploration and limits further improvement. Existing methods address this issue either through algorithm-level interventions, such as reward modification and entropy/KL regularization, or through token-level reweighting. We investigate a complementary perspective: entropy collapse can also be mitigated by changing which generated rollouts contribute to policy updates. Under the same sampling budget, not all rollouts contribute positively to an update, and selectively excluding some can improve learning. To address this, we propose GRPODropout: before the sta
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית