כתבה
arXiv cs.AI ·
Learn from the Gap: Differential-Aware Advantage Pruning with Adaptive Rollout Sampling for GRPO
תקציר מקורי באנגליתarXiv:2609.36932v1 Announce Type: new Abstract: Recently, Group Relative Policy Optimization (GRPO) and its variants have been developed for policy optimization and demonstrated notable performance gains. However, these methods usually incur substantial computational overhead due to per-question multi-rollout sampling and repeated per-token probability evaluation across rollouts. Furthermore, low-information or highly homogeneous trajectories can degrade downstream learning signal efficiency, hindering model optimization and limiting final performance. To address these issues, we propose FastRL, a novel plug-and-play reinforcement learning framework that simultaneously improves training efficiency and the effectiveness of policy learning. Specifically, 1) We introduce an advantage-aware pr
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית