כתבה
arXiv cs.AI ·
AgentLens: Adaptive Visual Modalities for Human-Agent Interaction in Mobile GUI Agents
תקציר מקורי באנגליתarXiv:2604.20279v3 Announce Type: replace-cross Abstract: Mobile GUI agents can automate smartphone tasks by interacting directly with app interfaces, but how they should communicate with users during execution remains underexplored. Existing systems rely on two extremes: foreground execution, which maximizes transparency but prevents multitasking, and background execution, which supports multitasking but provides little visual awareness. Through iterative formative studies, we found that users prefer a hybrid model with just-in-time visual interaction, but the most effective visualization modality depends on the task. Motivated by this, we present AgentLens, a mobile GUI agent that adaptively uses three visual modalities during human-agent interaction: Full UI, Partial UI, and GenUI. Agen
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית