יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

LSR-Ben: בנק הערכה לוגי ומדעי לבדיקת דגמי שכר תהליכי

LSR-Ben: A Logical and Scientific Reasoning Benchmark for Evaluating Process Reward Models
בנק הערכה חדש לבדיקת דגמי שכר תהליכי לשכילות ולמדעיות. הבנק, LSR-Ben, נועד לבדוק את יכולת הדגמים לזהות שגיאות בתהליכים. נמצא שהדגמים הקיימים חסרים ביכולתם לזהות שגיאות בתהליכים מעבר למתמטיקה.
תקציר מקורי באנגליתarXiv:2605.01203v3 Announce Type: replace Abstract: Currently, process reward models (PRMs) have exhibited remarkable potential for test-time scaling. Since large language models (LLMs) regularly generate flawed intermediate reasoning steps when tackling a broad spectrum of reasoning and decision-making tasks, PRMs are required to possess capabilities for detecting process-level errors in real-world scenarios. However, existing benchmarks primarily focus on mathematical reasoning, thereby failing to comprehensively evaluate the error detection ability of PRMs across diverse reasoning scenarios. To mitigate this gap, we introduce LSR-Ben, a process-level benchmark specifically designed for assessing PRM's performance across two primary reasoning domains (scientific and logical reasoning) an
קרא במקור המקורי