כתבה
arXiv cs.CL ·
Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech
תקציר מקורי באנגליתarXiv:2609.03502v1 Announce Type: new Abstract: In low-resource settings, deploying TTS typically requires choosing between a large voice-cloning model with costly inference or a compact fixed-voice system that requires a speaker-specific corpus. We study a third route: using a large voice-cloning model as a programmable data source to turn a short voice reference (e.g., 15 seconds) into a compact fixed-voice student trained entirely on synthetic speech. This setting makes pipeline design consequential: teacher errors become training targets, while filtering failed generations can reduce coverage of difficult texts. Thai further introduces challenges from ambiguous word boundaries, lexical tone, names and loanwords, numeric verbalization, and Thai-English code-switching. We study how text
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית