כתבה
arXiv cs.AI ·
לא נלמד, או לא נמדד? על רגישות הציונים לקושי ב-RLVR
Unlearnable, or Unmeasured? On the Reliability of Difficulty Labels in RLVR
במאמר זה נבחן את התופעה של קושי שאינו נלמד ב-RLVR. נמצא כי הקושי המוגדר נמצא לא רגיש, והסיבה לכך היא שהציונים נמדדים על סמך ספירה מוגבלת של תשובות. נציג פרקטיקה לקביעת רגישות הציונים.
תקציר מקורי באנגליתarXiv:2609.40115v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has become an important approach for improving reasoning during post-training. Recent work suggests that some difficult prompts remain resistant to learning even when they occasionally produce correct solutions. We revisit this unlearnability phenomenon and find that the affected prompts do improve, at roughly one third of the learnable rate, while the difficulty-defined set used to study them is much less reproducible than expected. These difficulty labels are estimated from a limited number of sampled responses. Combining them across seeds can further change which prompts are selected instead of simply reducing measurement noise. We develop a sampling-based framework for quantifying this
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית