כתבה
arXiv cs.CL ·
COT-TTS: דיבור מודע להקשר עם תהליך רצוני
COT-TTS: Audio Context-Aware Text-to-Speech with Chain-of-Thought Reasoning
COT-TTS הוא מערכת דיבור מודעת להקשר המשתמשת בתהליך רצוני. המערכת מסוגלת להבין הקשר קונברסיאלי, להסיק מסקנות ברורות ולייצר דיבור עם גוון מוגדר. המחקר כולל מערכת אימון גדולה ובדיקות מוצלחות.
תקציר מקורי באנגליתarXiv:2609.22697v2 Announce Type: replace Abstract: Recently, text-to-speech systems have made significant progress in speech expressiveness and controllability. However, the speaking style of generated speech typically relies on clear user-specified instructions. In natural conversations, speaking style should be naturally inferred from the preceding conversational context. Therefore, we propose COT-TTS, a context-aware, reasoning-based text-to-speech task. Given historical conversation audio, target text, and a reference speech, the system should comprehend the conversational context, infer an explicit intermediate reasoning, and finally synthesize the target speech with the specified timbre. To support this task, we constructed a large-scale bilingual conversational speech dataset compr
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית