כתבה
arXiv cs.LG ·
CompMat-Bench: Benchmarking AI Agents for Computational Materials Science
תקציר מקורי באנגליתarXiv:2610.00636v1 Announce Type: cross Abstract: Evaluating AI agents on scientific research tasks is constrained by the time and resources required for the underlying experiments or calculations. In computational materials research, repeating the same expensive simulations across agents and trials can make evaluation impractical. We introduce CompMat-Bench, a benchmark of 94 tasks derived from recently published computational materials studies, each asking agents to complete a step toward achieving the study's scientific goal. We reproduce the research steps in advance and assess agents on preparing inputs and analyzing outputs for expensive simulations, so expensive simulations can be avoided during evaluation. The reproduced inputs and results serve as ground truth for grading agents w
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית