כתבה
arXiv cs.CL ·
EmphTTS: an emphasis-control TTS with reinforcement learning
תקציר מקורי באנגליתarXiv:2609.27599v2 Announce Type: cross Abstract: Generating controllable and human-like emphasis remains an open challenge in text-to-speech, even when explicit emphasis control signals are provided in the text input, limiting the communicative accuracy of synthetic speech in real-world applications. Reinforcement learning has recently shown promise for post-training TTS systems to align with human preference, yet existing methods have not been applied to word-level prosodic control. We present EmphTTS, a non-autoregressive TTS system that applies Group Relative Policy Optimization (GRPO) to the duration predictor with an emphasis localization reward, enabling direct optimization for word-level emphasis. Evaluations show that EmphTTS achieves the best emphasis controllability and performs
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית