כתבה
arXiv cs.CL ·
בנק התזה של יישור
Robust Reasoning Benchmark
אנו מציגים את בנק התזה של יישור, פיילוט של 13 פרטורציות טקסטואליות סדרתיות שנלמדו על AIME 2024 ו-AIME 2025. הבנק מבדוק את יכולת הפתרון של 8 מודלי LLM מובילים, ומציג תוצאות מעניינות.
תקציר מקורי באנגליתarXiv:2604.08571v3 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) achieve high performance on standard mathematical benchmarks, their problem-solving abilities depend on the context and textual formatting. We introduce the Robust Reasoning Benchmark (RRB), a pipeline of 13 deterministic textual perturbations applied to AIME 2024 and AIME 2025. Evaluating 8 state-of-the-art models, we find that frontier models are largely resilient, with the notable exception of Claude, which categorically refuses many transformed prompts. Open-weights reasoning models exhibit a range of failure modes under structural noise (cognitive thrashing, tokenization breakdown, and reasoning collapse), with up to 54% average accuracy drops across perturbations and up to 100% on some. We fu
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית