כתבה
arXiv cs.CL ·
Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis
תקציר מקורי באנגליתarXiv:2607.24430v1 Announce Type: cross Abstract: Conversational Speech Synthesis is a fundamental component of human-computer interaction, aiming to generate contextually appropriate, expressive, and empathetic speech. However, facial expressions encode subtle and rich affective cues that are crucial for empathetic speech interaction, whereas existing approaches often overlook this important modality. In addition, the lack of large-scale natural conversational datasets with both speech and visual modalities also limits the development of visual affect understanding in conversational settings.To address these limitations, we propose FacialTalker, a facial-expression-aware CSS framework built upon a large language model backbone. To efficiently encode facial expressions, we propose AUTokeni
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית