כתבה
arXiv cs.CL ·
ConsensusBench: בנצ'מרק לצמתים מוסכמים
ConsensusBench: Benchmark of Consensus Nodes for LLM Reasoning via Outcome Reward Densifying
ConsensusBench הוא בנצ'מרק חדש לבדיקת יכולות ההיגיון של מודלים גדולים של שפה. הוא משתמש בצמתים מוסכמים כדי לספק אותות תהליך לתגובות למידה חיזוקית. המחקר הראה שיפור בביצועים לעומת שיטות קודמות.
תקציר מקורי באנגליתarXiv:2609.04648v1 Announce Type: new Abstract: Reinforcement learning (RL) has become one of the primary paradigms for reasoning enhancement of large language models (LLMs). In particular, Group Relative Policy Optimization (GRPO) and related algorithms have demonstrated strong performance with outcome-level rewards. However, these methods depend solely on the final answer, without feedback regarding which intermediate steps contribute to success or failure. As task complexity and reasoning trajectory length increase, such sparse final-answer rewards become increasingly insufficient. To address this limitation, we introduce ConsensusBench, a novel dataset designed to provide rule-based process-level signals. We posit that a correct final answer relies on a small set of intermediate conclu
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית