כתבה
arXiv cs.CL ·
Tacit-TTS: שיבוט קול מהיר ויעיל
Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Tacit-TTS הוא מערכת שיבוט קול מהירה ויעילה, שמאפשרת שיבוט קול ללא תלות בתעתיקים. המערכת משתמשת בגישה חדשה של יצירה לא-אוטורגרסיבית, ומאפשרת יצירה מהירה של קולות באיכות גבוהה.
תקציר מקורי באנגליתarXiv:2609.38658v1 Announce Type: cross Abstract: TTS systems with autoregressive semantic modeling have demonstrated strong zero-shot voice cloning performance and rich expressive variation, but their sequential decoding incurs substantial latency. Non-autoregressive alternatives offer much faster generation, yet often rely on more restrictive reference conditioning, such as requiring transcripts of the reference speech during inference. We present Tacit-TTS, an efficient transcript-free zero-shot voice cloning system distilled from IndexTTS2. Our model replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation, introduces training-free acoustic length estimation, and accelerates the flow-matching renderer through ReFlow distillation. Across two English an
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית