יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

Harbor Adapters ו-Harbor-Index: תשתית ומאגר מטא-נתונים מאומת לבדיקות גדולות של יצירתיות

Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation
Harbor Adapters ו-Harbor-Index הם תשתית ומאגר מטא-נתונים שמאפשרים בדיקות גדולות של יצירתיות. הם כוללים 82 תרגילים קשים ומגוונים, שנבחרו מתוך 29 בסיסי בדיקה. התשתית והמאגר נועדו לשפר את יכולת הבדיקה של יצירתיות, ולאפשר למדענים לבדוק את יצירתיותם של רבעים שונים.
תקציר מקורי באנגליתarXiv:2609.04298v2 Announce Type: replace-cross Abstract: Evaluating agents on the growing number of agentic benchmarks is challenging because they often require complex environments and agent integrations. We introduce Harbor Adapters, a unified evaluation infrastructure for agentic benchmarks. Our work makes three contributions. First, we develop benchmark adapters that port more than 80 benchmarks to evaluate arbitrary agents, and validate them through rigorous code review and parity experiments. Second, we conduct a large-scale evaluation of 8 models spanning capability tiers across 54 benchmarks; every model is run with Terminus-2 and with one of 3 native harnesses. This enables a broader analysis of agent capabilities and failure modes than was previously possible. Third, we introduc
קרא במקור המקורי