כתבה
arXiv cs.AI ·
PerfReasoning: כמה טובים ה-LLMs בהסבר יכולת הביצועים של התקרבות?
PerfReasoning: How Well Do LLMs Reason on Hardware Performance?
במבחן זה, ה-LLMs בודקים את יכולתן להסביר את יכולת הביצועים של התקרבות, וכן לייצר קוד למודלי ביצוע. המבחן כולל תכונות של גמיני ו-GPT-5.
תקציר מקורי באנגליתarXiv:2609.04476v3 Announce Type: replace Abstract: Performance modeling is central to hardware design and software optimization, yet constructing these models requires structured reasoning about computation, data reuse, storage, and movement. We introduce PerfReasoning, a benchmark that evaluates LLMs both as direct performance reasoners and as generators of analytical performance-model code. Given workload, architecture, and mapping specifications, models compare mappings and predict off-chip traffic and buffer requirements. The strongest closed-source models exceed 90% on reasoning-based Q&A, and the best open-weight model reaches 82.4%. However, model construction is substantially harder: while GPT-5.6 Sol exceeds 80% pass rate, all other model configurations average below 45% and vary
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית