יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

CompMat-Bench: בנצ'מרק לסוכנויות AI במדע חומרים

CompMat-Bench: Benchmarking AI Agents for Computational Materials Science
CompMat-Bench הוא בנצ'מרק לסוכנויות AI במדע חומרים. הוא מאפשר הערכה של סוכנויות AI על ידי 94 משימות מחקר. הבנצ'מרק מאפשר השוואה בין סוכנויות AI וניתוח סיבות כישלון.
תקציר מקורי באנגליתarXiv:2610.00636v1 Announce Type: new Abstract: Evaluating AI agents on scientific research tasks is constrained by the time and resources required for the underlying experiments or calculations. In computational materials research, repeating the same expensive simulations across agents and trials can make evaluation impractical. We introduce CompMat-Bench, a benchmark of 94 tasks derived from recently published computational materials studies, each asking agents to complete a step toward achieving the study's scientific goal. We reproduce the research steps in advance and assess agents on preparing inputs and analyzing outputs for expensive simulations, so expensive simulations can be avoided during evaluation. The reproduced inputs and results serve as ground truth for grading agents wit
קרא במקור המקורי