כתבה
arXiv cs.CL ·
תעתיק הגייה ללא אימון
Training-Free Pronunciation Transcription via Text-Constrained Acoustic Rescoring
חוקרים הציגו שיטה חדשה לתעתיק הגייה ללא אימון, המשלבת מידע לקסיקלי ואקוסטי. השיטה משתמשת בכלים G2P ו-S2P קיימים, ומשיגה דיוק גבוה יותר משיטות קודמות.
תקציר מקורי באנגליתarXiv:2609.30924v1 Announce Type: new Abstract: Accurate and efficient pronunciation transcription is essential for preparing text-to-speech training data at scale. Existing approaches have different limitations: grapheme-to-pronunciation (G2P) and speech-to-pronunciation (S2P) methods each capture only partial information, using only text or only speech, while speech-and-text-to-pronunciation (ST2P) methods use both but require costly pronunciation-annotated data. To address this problem, we propose a training-free ST2P pipeline that integrates both lexical and acoustic information at inference time. Lexical resources and G2P tools generate text-constrained candidates, and a left-to-right greedy search selects the best one using whole-sequence negative log-likelihoods from frozen pretrain
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית