כתבה
arXiv cs.AI ·
cua-speedrun: תקן תקין לבדיקת מהירות של סוכני חיפוש במחשב
cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents
cua-speedrun הוא תקן תקין שמטרתו לבדוק את מהירותם של סוכני חיפוש במחשב. התקן כולל תקני חיפוש ופלטפורמה תקינה שמאפשרת לבדוק את מהירותם של סוכנים רבים. התקן כולל גם תקני חיפוש ופלטפורמה תקינה שמאפשרת לבדוק את מהירותם של סוכנים רבים.
תקציר מקורי באנגליתarXiv:2609.40284v1 Announce Type: cross Abstract: Computer use agents (CUAs), which use graphical user interfaces (GUIs) to complete tasks on a computer, have recently surpassed human performance on many standard benchmarks, including difficult long-horizon tasks. Their capabilities are undoubtedly impressive, however, a key barrier to the widespread adoption and deployment of CUAs remains their speed and cost. Progress towards faster yet capable CUAs requires reliable evaluation of their speed, but many CUA benchmarks currently face a reproducibility crisis. Benchmarks are based on complex infrastructure with varying machine and container configurations that confound the evaluation of the execution speed of CUAs. Towards addressing this gap, we propose cua-speedrun, which introduces stand
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית