כתבה
arXiv cs.AI ·
ST-Bench: בנצ'מרק למערכות רב-סוכנים
ST-Bench: A Spatial-Temporal Benchmark for Multi-Agent System Generation on Scientific Research Tasks
ST-Bench הוא בנצ'מרק למערכות רב-סוכנים. הוא בודק האם מערכות רב-סוכנים משתפרות על סוכנים בודדים בנושאי ניתוח נתונים מדעיים. הבנצ'מרק כולל 100 משימות ניתוח נתונים מתחומי מדעי כדור הארץ.
תקציר מקורי באנגליתarXiv:2610.07763v1 Announce Type: new Abstract: The rapid progress of LLM-based multi-agent systems (MAS) has shown that they largely outperform single agents on coding, math, and QA tasks, where executable tests provide a binary success signal. Whether this advantage transfers to real scientific data analysis remains untested. We introduce ST-Bench, a benchmark designed to answer two questions: whether MAS outperform single agents on complex scientific data analysis tasks, and if so, by how much and at what additional cost. ST-Bench contains 100 data science tasks adapted from published Earth science studies across hydrology, agriculture, and wetland methane research, expanded into 2,067 queries grounded in additional published studies and validated by domain experts. Using ST-Bench, we e
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית