כתבה
arXiv cs.CL ·
EVOHARNESSBENCH: האם סוכנים יכולים לשמור קצב עם הרשת המתפתחת?
EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?
EVOHARNESSBENCH היא בסיסת-מבחן לבדיקת סוכנים תחת השפעת-קביעה מבוקרת של הרשת. היא כוללת 17 זרמי-רשת מבוססי-מערכת שנבנו באופן תקין מבסיסי-בדיקה של וריפיקטור, הכולל 802 תפקידים, 520 כלים, 42 כישורים ו-62 סוכנים.
תקציר מקורי באנגליתarXiv:2609.04280v1 Announce Type: cross Abstract: Modern LLM-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes what they observe and what they can do. In practice, this harness continually evolves as new capabilities are added. We introduce EVOHARNESSBENCH, a benchmark for evaluating agents under controlled harness evolution across three axes (tools, skills, and agents). Unlike existing continual-learning benchmarks for agents, which typically place non-stationarity (i.e., what changes over time) in the task stream while keeping the harness fixed, EVOHARNESSBENCH places non-stationarity in the externally supplied harness itself. It contains 17 multi-stage harness streams constructed deterministically from verifier-based benchmarks, comprisi
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית