כתבה
arXiv cs.AI ·
OpenProblemBench: בנק אתגרים לאינטליגנציה מלאכותית
OpenProblemBench: Benchmarking AI on Open Problems in the Foundational Theoretical Sciences
OpenProblemBench הוא בנק אתגרים לבדיקת יכולות האינטליגנציה המלאכותית בפתרון בעיות מדעיות בלתי פתורות. הבנק מכיל 82 בעיות מתוך ספרות המתמטיקה והפיזיקה. דגם GPT-6-Astra השיג את שיעור הפתרון הגבוה ביותר.
תקציר מקורי באנגליתarXiv:2610.11118v1 Announce Type: new Abstract: The next frontier for artificial general intelligence is tackling unresolved scientific problems, calling for benchmarks that assess progress beyond established knowledge. We introduce OpenProblemBench, a benchmark of 82 unresolved problems drawn from the mathematics and theoretical physics literature. Each problem supplies the research context, assumptions, and prior progress needed to investigate the question. We select problems whose proposed solutions admit comparatively clear checks of their decisive mathematical or computational claims. Four evaluator models independently assess the correctness, completeness, and degree of progress of each submission without reference solutions. Across seven evaluated configurations, GPT-6-Astra achieve
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית