יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

GRPODropout: Less is More for Online Reinforcement Learning Rollouts

תקציר מקורי באנגליתarXiv:2610.11854v1 Announce Type: new Abstract: Reinforcement learning (RL) methods such as GRPO substantially improve large language model reasoning but often suffer from policy entropy collapse: the loss of sampling diversity weakens exploration and limits further improvement. Existing methods address this issue either through algorithm-level interventions, such as reward modification and entropy/KL regularization, or through token-level reweighting. We investigate a complementary perspective: entropy collapse can also be mitigated by changing which generated rollouts contribute to policy updates. Under the same sampling budget, not all rollouts contribute positively to an update, and selectively excluding some can improve learning. To address this, we propose GRPODropout: before the sta
קרא במקור המקורי