יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

עבר לעבר: תכנון סגנון עם פרסום דיבור לטקסט-לדיבור

Beyond Speech Captions: Speech-Rewarded Style Planning for Conversational Text-to-Speech
תכנון סגנון עם פרסום דיבור לטקסט-לדיבור מאפשר טקסט-לדיבור תקשורתי. המאמר מציג פתרון חדש שמטפל בסגנון הדיבור ומשפר את תפקוד הטקסט-לדיבור.
תקציר מקורי באנגליתarXiv:2610.11461v1 Announce Type: cross Abstract: Natural-language style descriptions provide an interpretable interface between large language models (LLMs) and controllable text-to-speech (TTS). However, using descriptions as pseudo-labels compresses target acoustics into text, and descriptive fidelity need not imply effective control of a particular synthesizer. We empirically show that speech-text alignment only weakly predicts downstream acoustic similarity among candidate instructions for the same utterance. We therefore propose Speech-Rewarded Style Planning (SRSP), which trains a text-based style planner through a frozen downstream TTS model. Given dialogue history and response text, the planner generates candidate instructions and is optimized with group-relative policy optimizati
קרא במקור המקורי