כתבה
arXiv cs.CL ·
Diversifying RLVR Rollouts via First-Token Exploration
תקציר מקורי באנגליתarXiv:2605.28295v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) trains reasoning models without labeled trajectories, using groups of verifier-scored rollouts to explore alternative reasoning paths. Limited rollout diversity is a central bottleneck, typically addressed through adjustments to temperature, prefixes, or rollout selection. We identify the first token of the response as a structurally distinct target for diversification, largely overlooked in prior work. We find that the first-token distribution is sharply concentrated and only weakly related to downstream correctness, as lower-probability candidates can yield similarly accurate responses. Diversifying the first token can therefore broaden the reasoning paths explored within each
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית