כתבה
arXiv cs.LG ·
רצועה הגנטי: תפיסה של משימות אגנטיות
Running the Gauntlet: Challenging Agentic Tasks
במאמר זה, נוצרה רצועה הגנטית של משימות אגנטיות, כדי לבדוק את יכולות האגנטים. הרצועה כוללת 135 משימות, והאגנטים המובילים עדיין לא הצליחו להשיג יכולות אנושיות.
תקציר מקורי באנגליתarXiv:2606.14397v4 Announce Type: replace Abstract: As agentic systems continue to evolve and are widely deployed in real-world scenarios, there is a growing demand to faithfully evaluate their capabilities. However, current benchmarks are typically built on popular applications with relatively simple tasks and focus on a narrow set of capabilities while overlooking broader dimensions, resulting in saturated performance on modern agents and failing to probe their limitations. To this end, we introduce GauntletBench, a web-based benchmark for evaluating agent generalisation in challenging scenarios, focusing on three underexplored capabilities (temporal perception, graphical understanding, and 3D reasoning), across five less-covered professional applications (Video Editor, Workflow Builder,
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית