כתבה
arXiv cs.LG ·
ClosureBench: תקן בנייה למחשבים לוגיים
ClosureBench: A Constructive Benchmark for Compositional Graph Reasoning
נציגים חדשים למחשבים לוגיים, כגון GPT-4.1 ו-Gemini 2.5, נבחנים ב-ClosureBench, תקן בנייה למחשבים לוגיים. התקן נועד לבדוק את יכולתם של המודלים לבצע סיבוכיות רב-שלבית.
תקציר מקורי באנגליתarXiv:2608.18242v2 Announce Type: replace Abstract: Large language models fail on multi-step compositional reasoning, but measuring that failure is hard, because new models are trained on the benchmarks used to evaluate them. A fixed test set becomes a memorisation check soon after release. Constructive benchmarks avoid this by generating instances on demand. We introduce ClosureBench, a constructive benchmark for graph-relational logical reasoning. Each task is built from explicit primitives (reachability, degree, set operations, connectivity, aggregation), and its reference answer is computed by executing code that implements that logic exactly. Ground truth is therefore verified, and the supply of fresh instances is unlimited. The benchmark spans 26 task categories at three compositiona
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית