כתבה
arXiv cs.CL ·
מה MMLU באמת מודד?
What Does MMLU Actually Measure? A Psychometric Audit of Difficulty Structure in Aggregate Benchmark Scores
MMLU הוא בנקאות נפוץ לכיול יכולות AI כלליות, אך מחקר זה מראה כי הוא בעיקר מודד יכולת שליפה עובדתית ולא יכולת תהלichה. המחקר מציע כי יש לדווח בנפרד על יכולות אלו.
תקציר מקורי באנגליתarXiv:2609.09372v1 Announce Type: cross Abstract: Although MMLU is widely adopted as a benchmark for calibrating general AI capabilities, we psychometrically demonstrate that its aggregate score primarily evaluates a model's factual retrieval capacity rather than its reasoning ability. By calibrating item difficulty for 1,000 open-weights language models over 14,042 MMLU test items using Item Response Theory, we show that evaluating both abilities via a single test is inherently flawed. Difficulty is then regressed on a deterministic, text-extractable framework of structural complexity. Applying a joint Wald test with subject-clustered covariances demonstrates that the MMLU conflates fundamentally separable constructs. The mapping from structural complexity to difficulty is not invariant a
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית