כתבה
arXiv cs.AI ·
AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use
תקציר מקורי באנגליתarXiv:2607.20536v1 Announce Type: new Abstract: Tool-use agents that address day-to-day digital tasks such as ordering groceries must not only operate applications, but also interact with the user, e.g., to ask clarification questions, prompt for confirmation, and inform the user when the instruction is infeasible. However, current benchmarks for evaluating agent-user interactions do not capture the diversity of such interactions. Further, they operate in small environments with few, often non-state-changing, APIs. To address this gap, we introduce AppWorld-UL, a ``user-in-the-loop'' benchmark of 516 challenging tasks requiring diverse agent-user interactions. Building upon the AppWorld framework with 9 popular simulated apps like Amazon and Spotify, we systematically modify original tasks
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית