יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

אגנטים: מערכות, ולא מודלים: חשיבה מחדש על בדיקת אגנטים

Agents Are Systems, Not Models: Rethinking Agentic Evaluation
במאמר זה, חוקרים חושבים מחדש על בדיקת אגנטים. הם מציעים לבדוק את האגנט כמערכת ניתנת להגדרה, ולא רק כמודל. הם מציגים נתונים חדשים ומציעים שיטות חדשות לבדיקת אגנטים.
תקציר מקורי באנגליתarXiv:2610.01618v1 Announce Type: new Abstract: Agent evaluations increasingly go beyond a single success rate, reporting metrics such as cost, consistency, and robustness. Yet they typically treat the agent itself as fixed. In practice, an agent is a configurable system: users decide what to tell it, how long to let it run, and which model to use, and each of these choices can change how well and how consistently it performs. We study these choices on a new benchmark of four scientific tasks, where a coding agent must find and correctly operate a published specialist model. We investigate five parts of the agent's configuration: task information, reasoning, self-verification, time budget, and backbone model. We find substantial run-to-run variability, with approximately 54% of the outcome
קרא במקור המקורי