כתבה
arXiv cs.CL ·
cua-speedrun: בנצ'מרקים למהירות של סוכנים
cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents
cua-speedrun הוא פרויקט לבנצ'מרקים מהירות ויעילות של סוכנים המשתמשים בממשקי משתמש גרפיים. הוא מציע תשתית וקבוצת משימות סטנדרטיות, ומאפשר השוואה בין סוכנים שונים. הפרויקט כולל ניתוח של השפעת מאמץ היגיון, ערכות סוכנים ועיכובי סביבה על ביצועים.
תקציר מקורי באנגליתarXiv:2609.40284v1 Announce Type: cross Abstract: Computer use agents (CUAs), which use graphical user interfaces (GUIs) to complete tasks on a computer, have recently surpassed human performance on many standard benchmarks, including difficult long-horizon tasks. Their capabilities are undoubtedly impressive, however, a key barrier to the widespread adoption and deployment of CUAs remains their speed and cost. Progress towards faster yet capable CUAs requires reliable evaluation of their speed, but many CUA benchmarks currently face a reproducibility crisis. Benchmarks are based on complex infrastructure with varying machine and container configurations that confound the evaluation of the execution speed of CUAs. Towards addressing this gap, we propose cua-speedrun, which introduces stand
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית