כתבה
arXiv cs.AI ·
MGSM-Pro: אסטרטגיה פשוטה לבחינת נימוק מתמטי רב-לשוני
MGSM-Pro: A Simple Strategy for Robust Multilingual Mathematical Reasoning Evaluation
MGSM-Pro הוא מאגר נתונים לבחינת נימוק מתמטי רב-לשוני. הוא מציע חמש גרסאות לכל שאלה, עם שינויים בשמות, ספרות והקשר. הבחינה נערכה בתשע שפות והראתה שמודלים רבים אינם עמידים לשינויים בספרות. מודלים כגון GPT-OSS 120B ו-DeepSeek v3 הראו עמידות חזקה יותר.
תקציר מקורי באנגליתarXiv:2601.21225v5 Announce Type: replace-cross Abstract: Large language models have made substantial progress in mathematical reasoning. However, benchmark development for multilingual evaluation has lagged behind English in both difficulty and recency. Recently, GSM-Symbolic showed a strong evidence of high variance when models are evaluated on different instantiations of the same question; however, the evaluation was conducted only in English. In this paper, we introduce MGSM-Pro, an extension of MGSM dataset with GSM-Symbolic approach. Our dataset provides five instantiations per MGSM question by varying names, digits and irrelevant context. Evaluations across nine languages reveal that many low-resource languages suffer large performance drops when tested on digit instantiations diffe
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית