כתבה
arXiv cs.LG ·
Conversational Voice Aesthetic Model with Reinforcement Learning from Human Listeners
תקציר מקורי באנגליתarXiv:2610.10868v1 Announce Type: cross Abstract: We introduce Conversational Voice Aesthetic Model, a speech large language model for describing the voice aesthetics of real or synthetic speech responses in natural conversational contexts. Given a context and a response speech, CVAM describes salient moments that characterize the voice and predicts nine categorical attributes spanning gender, pitch, pacing, emotion, and delivery. The key challenge lies in perceptual fields such as emotion and delivery, which are inherently subjective and lack definitive ground truth. Therefore, we collect ~10 human annotations for each of 3k real and synthetic responses derived from the CANDOR corpus. CVAM is supervised finetuned on synthesized aesthetic descriptions and labels, then optimized with Group
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית