יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

MGSM-Pro: סטרטגיה פשוטה לבדיקת תיקון חזק של חשבון מתמטי בשפות רבות

MGSM-Pro: A Simple Strategy for Robust Multilingual Mathematical Reasoning Evaluation
נחשפה סטרטגיה חדשה לבדיקת תיקון חזק של חשבון מתמטי בשפות רבות. המחקר כולל חידושים ב-GSM-Symbolic ו-MGSM-Pro. נמצא כי דגמי LLM חסרים חזקות בשפות נמוכות-משאבים. נמצא כי דגמי Gemini 2.5 Flash ו-GPT-4.1 פחות חסינים לדיגיטים, בעוד Gemini 3.0 Pro יותר חסינים. דגמי GPT-OSS 120B ו-DeepSeek v3 חסינים יותר.
תקציר מקורי באנגליתarXiv:2601.21225v4 Announce Type: replace-cross Abstract: Large language models have made substantial progress in mathematical reasoning. However, benchmark development for multilingual evaluation has lagged behind English in both difficulty and recency. Recently, GSM-Symbolic showed a strong evidence of high variance when models are evaluated on different instantiations of the same question; however, the evaluation was conducted only in English. In this paper, we introduce MGSM-Pro, an extension of MGSM dataset with GSM-Symbolic approach. Our dataset provides five instantiations per MGSM question by varying names, digits and irrelevant context. Evaluations across nine languages reveal that many low-resource languages suffer large performance drops when tested on digit instantiations diffe
קרא במקור המקורי