כתבה
arXiv cs.AI ·
APPSim-Bench: Bridging Real-world Apps and Reproducible Evaluation for Mobile GUI Agents
תקציר מקורי באנגליתarXiv:2609.07712v1 Announce Type: new Abstract: Mobile GUI agents can execute tasks from natural-language instructions, but their evaluation remains difficult to make both realistic and reproducible. Existing benchmarks typically trade off these goals: simplified apps lack real-world mobile complexity, whereas live commercial apps introduce uncontrolled variation from recommendations, advertisements, accounts, and changing content. We propose AppSim-Bench, which addresses this trade-off through controllable simulated apps that preserve task-relevant interaction logic while supporting deterministic evaluation. Built through a coding-agent-assisted and human-verified workflow, it contains 557 tasks across 17 high-frequency Chinese and English apps. Its controllable backend data and outcome-b
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית