כתבה
arXiv cs.AI ·
ElderBench: Benchmarking Autonomous Mobile Agents for Older Adults
תקציר מקורי באנגליתarXiv:2609.04850v1 Announce Type: new Abstract: While autonomous mobile agents hold great potential for assisting older adults with smartphone usage, existing GUI benchmarks mainly rely on explicit, goal-oriented instructions and rarely capture the naturally occurring language patterns of older users, such as indirect speech, referential ambiguity, and under-specified requests. This mismatch between benchmark instructions and real-world elderly interactions may hinder reliable agent deployment. To address this gap, we present ElderBench, the first benchmark for evaluating mobile GUI agents in authentic elderly-oriented scenarios. ElderBench is constructed from 249 naturally elicited smartphone tasks collected from older adults across 20 applications. We first characterize the linguistic di
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית