יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

UtterTune: עריכה ובקרה של הגייה ב-TTS רב-לשוני

UtterTune: LoRA-Based Target-Language Pronunciation Edit and Control in Multilingual Text-to-Speech
UtterTune היא שיטה קלה לשיפור בקרת ההגייה במערכות TTS רב-לשוניות. היא מאפשרת שליטה על ההגייה ברמת הפונמה, בעודה שומרת על טבעיות ודמיון לדובר בהגדרת zero-shot.
תקציר מקורי באנגליתarXiv:2508.09767v4 Announce Type: replace-cross Abstract: We propose UtterTune, a lightweight method for adapting a multilingual text-to-speech (TTS) system built on a large language model (LLM). It improves control of pronunciation in the target language while preserving performance in the others. Although LLM architectures have enabled TTS models to achieve remarkable naturalness, accurately modeling grapheme-to-phoneme (G2P) mapping and prosody remains challenging, especially when the model omits an explicit G2P module and directly processes minimally encoded text (e.g., byte-pair encoding). UtterTune leverages low-rank adaptation to enable the control of segmental pronunciation and pitch accent at the phoneme level for Japanese speech, the target language in this paper, while maintaini
קרא במקור המקורי