כתבה
arXiv cs.CL ·
When Trivia Is Not Trivial: Everyday Knowledge Failures in Multilingual LLMs
תקציר מקורי באנגליתarXiv:2607.21445v1 Announce Type: new Abstract: Quiz rooms, trivia nights, and quiz shows challenge human knowledge across a wide range of topics, from canonical facts to everyday culture. In this paper, we examine whether large language models (LLMs) can perform competitively in such settings, using quiz-style questions to test them on both common and niche topics. We introduce TriviaRoomQA, a multilingual benchmark designed to evaluate everyday, culturally grounded, and long-tail knowledge across 288 topics. The benchmark contains 3,300 parallel multiple-choice questions in six European languages and additional 5,340 French-only questions for a more fine-grained case study. We evaluate 30 open-weight LLMs from European, Asian, and North American providers, covering models from 7 to 70B p
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית