כתבה
arXiv cs.CL ·
Exploring Multimodal Turn-Taking Cues in Face-to-Face Conversation using Voice Activity Projection
תקציר מקורי באנגליתarXiv:2609.14666v1 Announce Type: cross Abstract: Turn-taking is a fundamental component of spoken interaction, and while humans naturally rely on both verbal and non-verbal signals, dialogue systems usually depend on audio cues alone. This paper investigates whether visual features from face-to-face conversations can enhance turn-taking prediction beyond what is achievable from audio-only. We extend the Voice Activity Projection (VAP) model, a self-supervised transformer-based model for predicting future voice activity, by incorporating visual features extracted from the large-scale Meta Seamless Interaction dataset of dyadic face-to-face conversations. The visual features include gaze direction, head movement, body and hand pose, and facial action units (FAU). For incorporating the visua
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית