יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

חקירה בנושא רובוסטנסיית LLMs בתחום המתמטיקה: השוואה עם תרגום-שינוי של בעיות מתמטיות מתקדמות

An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
חקירה בנושא רובוסטנסיית LLMs בתחום המתמטיקה, כולל תרגום-שינוי של בעיות מתמטיות מתקדמות. המחקר כולל תרגום-שינוי של 5,255 בעיות מתמטיות, ובדיקה של 18 מודלי LLM שונים.
תקציר מקורי באנגליתarXiv:2508.08833v4 Announce Type: replace-cross Abstract: Frontier large language models (LLMs) now reach near-ceiling accuracy on standard mathematical-reasoning benchmarks and gold-medal-level performance at the International Mathematical Olympiad. As these benchmarks saturate and their items leak into training data, a high score no longer shows whether a model reasons robustly or which component of its reasoning fails. To evaluate reasoning while keeping results informative and failures diagnosable, we propose GAP (Generalisation-and-Perturbation), a methodology that automatically generates mathematically equivalent variants of existing mathematics problems at scale using two disjoint, interpretable transformations: (1) surface renames, probing the binding between identifiers and latent
קרא במקור המקורי