כתבה
arXiv cs.AI ·
בדיקת BAIM-LLM: תקן לבדיקת כישורי כלי תכנות חיים
Agentic BAIM-LLM Evaluation (ABLE): Benchmarking LLM Use of Protein Design Tools
במאמר זה, נבדיק כישורי LLMs בשימוש בכלי תכנות חיים לתכנון חלבונים. נבדוק 15 מודלי LLM חדשים ונשווה את התוצאות לבסיס תקן אנושי.
תקציר מקורי באנגליתarXiv:2609.05818v1 Announce Type: new Abstract: We introduce ABLE, a benchmark for evaluating LLM agents' ability to use biological AI models (BAIMs), such as ProteinMPNN and AlphaFold3, in dual-use protein design workflows. ABLE assesses agent performance through a set of tasks spanning structure retrieval, sequence generation, and design validation. We evaluate 15 frontier models and find that seven refuse all tasks, while the remaining models exhibit substantial performance differences. Claude Sonnet 4 and Gemini 3 Pro achieve the highest scores across information retrieval, tool selection, and tool use. We further compare model performance on a subset of tasks against an expert human baseline. Our results suggest that current LLMs can substantially lower barriers to protein design, but
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית