כתבה
arXiv cs.CL ·
HyperLogic: מבחן סקלאבלי להסקה לוגית סינית
HyperLogic: A Hard, Forward-Authored Chinese Logical Reasoning Benchmark with Execution-Derived Answers
מבחן חדש להסקה לוגית בסינית, כולל שמות מודלים/כלים/חברות.
תקציר מקורי באנגליתarXiv:2605.19597v2 Announce Type: replace Abstract: Existing logic benchmarks primarily measure models' ability to answer reasoning questions directly. Scalable benchmarks often generate text from formal structures, which makes answers easy to compute but fixes the formalization before the problem is written. Forward construction preserves the challenge of finding a faithful formalization, yet makes difficulty and answer reliability harder to control. We introduce HyperLogic, a forward-construction pipeline that separates problem authoring from answer generation. A multi-agent workflow hardens undergraduate-authored Chinese seeds without solving them; two agents from different model families independently translate each finished item into executable finite-domain models; their encodings an
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית