יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

מדידת התקדמות בתפיסה לקבלת אינספירציה מתמטית עם וריפיקציה אוטומטית

Measuring Progress in Reasoning Toward Mathematical Discovery with Automatic Verification
אנו מדדים את התקדמות ה-AI בקבלת אינספירציה מתמטית. המחקר כולל 113 בעיות מתמטיות שאינן פתורות, ומציג כלי חדש לבדיקת התקדמות ה-AI. נמצאו שישה פתרונות חדשים לבעיות מתמטיות שהצליחו לפתור.
תקציר מקורי באנגליתarXiv:2603.15617v2 Announce Type: replace Abstract: Can AI make progress on important, unsolved mathematical problems? Large language models are now capable of sophisticated mathematical and scientific reasoning, but whether they can perform novel research is still widely debated and underexplored. We introduce HorizonMath, a benchmark of 113 predominantly unsolved problems spanning eight domains in mathematics and the mathematical sciences, paired with an open-source evaluation framework for automated verification. Our benchmark targets the generator-verifier gap: problems where discovery is hard and requires meaningful mathematical insight, but verification is computationally straightforward. This contrasts with most existing research-level benchmarks, which instead rely on formal proof
קרא במקור המקורי