יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

PETSCAgent-Bench: מסגרת אבחונית לקוד מדעי מיוצר

An Agentic Evaluation Framework for AI-Generated Scientific Code in PETSc
PETSCAgent-Bench היא מסגרת לאבחון קוד מדעי מיוצר על ידי מודלים כמו LLMs. היא בודקת נכונות, ביצועים, איכות קוד והתאמה לספריות HPC. המחקר מראה שמודלים מתקדמים יוצרים קוד קריא ומובנה, אך מתקשים בנכונות ובהתאמה לספריות.
תקציר מקורי באנגליתarXiv:2603.15976v2 Announce Type: replace Abstract: While LLMs have accelerated scientific code generation, comprehensively evaluating generated code remains challenging. Many benchmarks emphasize functional correctness or task completion, which is insufficient for code built on production HPC libraries, where solver selection, API conventions, memory management, parallel awareness, and performance also matter. We introduce PETSCAgent-Bench, a multidimensional benchmark and agent-based framework for assessing whether AI-generated scientific code uses a production HPC library as an expert would. A tool-augmented evaluator compiles, executes, and measures code and combines deterministic checks with LLM-based assessments in a 14-evaluator pipeline spanning five categories: correctness, perfor
קרא במקור המקורי