יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

EvoHarnessBench: האם סוכני ה-LLM יכולים לשמור קצב עם הרשת המתפתחת?

EvoHarnessBench: Can Your Agents Keep Pace with an Evolving Harness?
EVOHARNESSBENCH היא בסיס אבחנה לבדיקת סוכני LLM תחת התפתחות רשת מאולתרת. הבסיס כולל 17 זרמי רשת מרוכזים ומוגדרים כך שהם כוללים 802 תפקידים, 520 כלים, 42 מיומנויות ו-62 סוכנים.
תקציר מקורי באנגליתarXiv:2609.04280v2 Announce Type: replace-cross Abstract: Modern LLM-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes what they observe and what they can do. In practice, this harness continually evolves as new capabilities are added. We introduce EVOHARNESSBENCH, a benchmark for evaluating agents under controlled harness evolution across three axes (tools, skills, and agents). Unlike existing continual-learning benchmarks for agents, which typically place non-stationarity (i.e., what changes over time) in the task stream while keeping the harness fixed, EVOHARNESSBENCH places non-stationarity in the externally supplied harness itself. It contains 17 multi-stage harness streams constructed deterministically from verifier-based benchmarks,
קרא במקור המקורי