כתבה
arXiv cs.CL ·
כשהטריוויה איננה טריוויה: כשלי ידע יומיומי ב-LLMs
When Trivia Is Not Trivial: Everyday Knowledge Failures in Multilingual LLMs
מודלי שפה גדולים נתקלים בקשיים בידע יומיומי, במיוחד בתרבות הפופולרית. המאמר עוסק בבדיקת יכולות של LLMs בשאלות טריוויה, ומציג נתונים על חוסר ידע חשוב בתחום זה.
תקציר מקורי באנגליתarXiv:2607.21445v2 Announce Type: replace Abstract: Quiz rooms, trivia nights, and quiz shows challenge human knowledge across a wide range of topics, from canonical facts to everyday culture. In this paper, we examine whether large language models (LLMs) can perform competitively in such settings, using quiz-style questions to test them on both common and niche topics. We introduce TriviaRoomQA, a multilingual benchmark designed to evaluate everyday, culturally grounded, and long-tail knowledge across 288 topics. The benchmark contains 3,300 parallel multiple-choice questions in six European languages and additional 5,340 French-only questions for a more fine-grained case study. We evaluate 30 open-weight LLMs from European, Asian, and North American providers, covering models from 7 to 7
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית