כתבה
arXiv cs.CL ·
הגברת TTS מבוסס פונמה ל-ASR
Scaling phoneme-based TTS augmentation for ASR: A unified pipeline and controlled study
מחקר זה מציג פייפליין מאוחד ל-TTS ל-ASR מבוסס פונמה. המחקר מראה כי אוגמנטציה אקראית משפרת את התוצאות על 11 סטים שונים. השיפור המרבי היה 19.3% לעומת בחירה אקראית.
תקציר מקורי באנגליתarXiv:2608.26697v2 Announce Type: replace Abstract: Synthetic speech offers scalable supervision for automatic speech recognition (ASR), but its benefit depends on text selection, reference speech, and augmentation scale. We present a phoneme-based TTS-to-ASR pipeline using a single TTS model jointly trained from scratch on Arabic, French, Italian, and Portuguese with the F5-TTS architecture and language-monolingual ASR systems cover 13 test sets. Across the synthesis-scale sweep, random augmentation improves over matched real-only continuation on 11 sets. In the selection comparison, PFGS improves over real-only training on 12 sets and over random selection on nine, with a maximum relative WER reduction of 19.3% against random selection. With target texts and synthesis counts fixed, refer
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית