כתבה
arXiv cs.AI ·
MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning
תקציר מקורי באנגליתarXiv:2610.02824v1 Announce Type: new Abstract: Rubric-based reinforcement learning extends reward-driven optimization to open-ended tasks by assigning partial credit to individual response requirements. However, rubric judges can assign a high criterion score even when the information or action it requires is absent from the response, a failure mode we term Vacuous Credit. Such awards persist after the required information is removed and can reverse the sign of a response's GRPO advantage. To address this problem, we introduce MetaRubric, which alternates evidence-aware policy optimization with response-guided rubric adaptation. We construct counterfactual counterparts by changing one task-relevant fact in each prompt. During policy optimization, credit is assigned only when the response
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית