כתבה
arXiv cs.AI ·
DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents
תקציר מקורי באנגליתarXiv:2609.06059v1 Announce Type: new Abstract: As large language models evolve from question-answering systems into general-purpose agents, evaluation must move beyond static answer correctness to assess multimodal perception, multi-step execution, tool use, and artifact delivery. However, existing benchmarks are often tied to specific task types, execution environments, or scoring protocols, limiting their comparability, interpretability, and reliability for deployment decisions. We introduce DAREBench (Deployment-Aware and Reliable Evaluation of Models as Agents), a benchmark designed to capture workload variation and support reliable agent evaluation. Built on a shared OpenClaw execution environment, DAREBench organizes 233 tasks selected and adapted from 22 source benchmarks into a $2
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית