כתבה
arXiv cs.AI ·
DriftTTS: ייצור קול מטקסט במספר מועט של צעדים
DriftTTS: Few-Step Text-to-Speech Without Distillation via Distribution-Matching Drift
DriftTTS הוא מודל ייצור קול מטקסט שאינו תלוי בהדרכה ממודל מורכב קודם. הוא משתמש במטרה של התאמת הפצה במרחב תכונות. תוצאותיו תואמות אלו של מודלים אחרים כגון Matcha-TTS.
תקציר מקורי באנגליתarXiv:2610.03390v1 Announce Type: cross Abstract: Few-step neural text-to-speech models often rely on short- ened diffusion or flow-matching schedules, or on distillation from pretrained multi-step teachers. To avoid these depen- dencies, we present DriftTTS, a few-step mel-spectrogram generator trained without a generative teacher, distillation, or adversarial discrimination. DriftTTS uses a distribution- matching drift objective in a mel-domain feature space defined by raw mels and a frozen masked-autoencoder encoder pretrained on the same LJSpeech training split. On-policy rollout trains the decoder on its own interme- diate states and supports inference up to the trained roll- out depth. On LJSpeech, DriftTTS at NFE=4 achieves 3.87 dB MCD and 3.7% WER, compared with 3.85 dB and 3.4% fo
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית