כתבה
arXiv cs.CL ·
UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding
תקציר מקורי באנגליתarXiv:2606.07167v2 Announce Type: replace Abstract: Meaningful multilingual evaluation must test models in the target language and educational context. Urdu, spoken by more than 230 million people, lacks a broad MMLU-style benchmark built from native educational sources. We introduce UrduMMLU, a benchmark of 26,389 Urdu MCQs across 26 subjects and five domains, collected from native Urdu MCQ banks and public examination PDFs. Unlike translation-based benchmarks, UrduMMLU combines academic subjects with content specific to Urdu and regional education. We label the exam-derived portion through dual human annotation with strict consensus filtering. We evaluate 30 LLMs under English and Urdu prompts, yielding 60 zero-shot evaluations, and further evaluate four open-source LLMs under multiple f
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית