כתבה
arXiv cs.AI ·
Talk2Agent: בנצ'מרק לממשקי קול לסוכנים טקסט
Talk2Agent: Benchmarking Voice Interfaces for Text Agents
Talk2Agent הוא בנצ'מרק להערכת יעילותם של ממשקי קול לסוכנים טקסט מבוססי LLM. הוא בונה גרסאות מדוברות של משימות ומעריך מגוון ממשקי קול, כולל מודלים ASR ו-LLM. המחקר מציג פריים-וורק להערכה תפקודית של ממשקי קול.
תקציר מקורי באנגליתarXiv:2609.38867v1 Announce Type: new Abstract: Large language model (LLM) computer-use agents are typically evaluated with clean written instructions, despite speech being an increasingly popular interface for interacting with such systems. Speech input introduces an additional failure point: transcription errors can alter task-critical entities, constraints, or targets before the agent begins reasoning, while conventional ASR metrics do not directly measure whether the information required for successful execution has been preserved. We introduce Talk2Agent, a benchmark for evaluating how effectively voice interfaces convey human-spoken instructions to LLM-based computer-use agents. Talk2Agent builds human-spoken versions of tasks from WildClawBench and OSWorld and evaluates a range of v
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית