יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

CoArena: בדיקת מערכות-סוכן ומערכות-סוכנים בזמן-אמת

CoArena: Evaluating Computer-Use and Multi-Agent Systems in Real Time
CoArena בודק מערכות-סוכן ומערכות-סוכנים בזמן-אמת. המערכת נותנת דירוגים ומדדים למערכות-סוכן, כולל גמיני.
תקציר מקורי באנגליתarXiv:2609.14239v1 Announce Type: new Abstract: Static benchmarks for computer-use agents fix a task set at release and score every system against it once. That makes them reproducible, and it lets them drift from what they should measure: a fixed task set ages, leaks into training corpora, and cannot follow how people actually use agents from week to week. CoArena measures use directly. Real users submit tasks; two systems, each a single model or a multi-agent pipeline behind the same tool interface, execute the same task concurrently in identical sandboxed desktops; users judge the two outcomes without knowing which system produced them; and a public leaderboard is refit from those judgments. The central contribution is a formal account of what makes such an evaluation real-time. We defi
קרא במקור המקורי