כתבה
arXiv cs.LG ·
RLVR: הקטנת גבולות ההיגיון
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion
RLVR (Reinforcement Learning with Verifiable Rewards) יכול לשפר דיוק בדגימה בודדת, אך לרעות בדגימה חוזרת. מחקר זה חוקר את התופעה הזו ומציע פתרון בשם Per-Problem Base Anchoring (PBA).
תקציר מקורי באנגליתarXiv:2607.20543v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) can improve one-sample accuracy while making a model worse under repeated sampling. We study this pass@k inversion: after training, the policy may solve fewer distinct problems than its base model at large $k$. The failure concentrates on boundary prompts, where the base model contains rare correct trajectories that are recoverable by sampling but too sparse to reliably appear in finite RLVR rollout groups. We argue that a two-mode account explains this as an absence-of-evidence failure: rare correct trajectories may disappear before RLVR samples and reinforces them often enough. The main contribution is this diagnostic and mechanistic framing. Per-Problem Base Anchoring (PBA) is a deliber
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית