כתבה
arXiv cs.LG ·
אופטימיזציה של פוליציה בעזרת קרדיטים-מאושרים חצי-סובטקטיביים
Semifactual Credit-Augmented Policy Optimization
אופטימיזציה של פוליציה בעזרת קרדיטים-מאושרים חצי-סובטקטיביים, כדי לשפר את ההסברה והכללה ב-RLVR.
תקציר מקורי באנגליתarXiv:2609.40360v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has improved the reasoning capabilities of large language models (LLMs), yet their predictions remain sensitive to task-irrelevant prompt features. We investigate this sensitivity through semifactual prompt interventions that preserve the underlying problem and its answer. Our analysis reveals substantial variation in token-level sensitivity and shows that suppressing high-drift token candidates during decoding improves reasoning accuracy without updating model weights. These findings highlight a limitation of Group Relative Policy Optimization (GRPO), which assigns the same outcome-derived advantage to every response token and may reinforce potential spurious dependence alongside useful r
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית