כתבה
arXiv cs.LG ·
QuArch: בנק המבחן לבדיקת יכולות התקשורת של LLM באדריכלות מחשבים
QuArch: A Benchmark for Evaluating LLM Reasoning in Computer Architecture
QuArch - בנק המבחן הראשון לבדיקת יכולות התקשורת של LLM באדריכלות מחשבים. הבנק הזה כולל 2,671 שאלות-תשובות מומחים-מאומתות, ומציג חוסרים ביכולות התקשורת של LLM באדריכלות מחשבים.
תקציר מקורי באנגליתarXiv:2510.22087v2 Announce Type: replace-cross Abstract: The field of computer architecture, which bridges high-level software abstractions and low-level hardware implementations, remains absent from current large language model (LLM) evaluations. To this end, we present QuArch (pronounced 'quark'), the first benchmark designed to facilitate the development and evaluation of LLM knowledge and reasoning capabilities specifically in computer architecture. QuArch v1.0 provides a comprehensive collection of 2,671 expert-validated question-answer (QA) pairs covering various aspects of computer architecture, including processor design, memory systems, and interconnection networks. Our evaluation reveals that while frontier models possess domain-specific knowledge, they struggle with skills that
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית