כתבה
arXiv cs.CL ·
VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents
תקציר מקורי באנגליתarXiv:2605.30256v2 Announce Type: replace-cross Abstract: Natural human conversation is full-duplex and audio-visual: people simultaneously speak and listen while continuously interpreting and producing nonverbal cues, such as nods, smiles, and gestures. To support successful human-agent interaction, agents must model full-duplex audiovisual conversation; however, existing full-duplex benchmarks evaluate only speech. In this work, we present VideoFDB, the first benchmark to evaluate full-duplex audio-visual-to-audio-visual (AV2AV) conversational agents. VideoFDB contributes (i) 237 dyadic clips spanning 11 nonverbal conversational dynamics from real-world video calls, (ii) a taxonomy separating perception from generation behaviors, and (iii) a rubric-based LM-as-judge evaluation framework
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית