כתבה
arXiv cs.CL ·
LSR-Ben: A Logical and Scientific Reasoning Benchmark for Evaluating Process Reward Models
תקציר מקורי באנגליתarXiv:2605.01203v3 Announce Type: replace-cross Abstract: Currently, process reward models (PRMs) have exhibited remarkable potential for test-time scaling. Since large language models (LLMs) regularly generate flawed intermediate reasoning steps when tackling a broad spectrum of reasoning and decision-making tasks, PRMs are required to possess capabilities for detecting process-level errors in real-world scenarios. However, existing benchmarks primarily focus on mathematical reasoning, thereby failing to comprehensively evaluate the error detection ability of PRMs across diverse reasoning scenarios. To mitigate this gap, we introduce LSR-Ben, a process-level benchmark specifically designed for assessing PRM's performance across two primary reasoning domains (scientific and logical reasoni
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית