יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

בדיקת עמידות של LLMs בנימוק מתמטי

An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
חוקרים בדיקת עמידות של מודלים גדולים של שפה (LLMs) בנימוק מתמטי. הם מציגים שיטה חדשה ליצירת נוסחאות מתמטיות שקולות, GAP, ובודקים 18 מודלים מסחריים ופתוחים. התוצאות מראות ירידה בדיוק עם שינויים בנוסחאות.
תקציר מקורי באנגליתarXiv:2508.08833v4 Announce Type: replace Abstract: Frontier large language models (LLMs) now reach near-ceiling accuracy on standard mathematical-reasoning benchmarks and gold-medal-level performance at the International Mathematical Olympiad. As these benchmarks saturate and their items leak into training data, a high score no longer shows whether a model reasons robustly or which component of its reasoning fails. To evaluate reasoning while keeping results informative and failures diagnosable, we propose GAP (Generalisation-and-Perturbation), a methodology that automatically generates mathematically equivalent variants of existing mathematics problems at scale using two disjoint, interpretable transformations: (1) surface renames, probing the binding between identifiers and latent varia
קרא במקור המקורי