כתבה
arXiv cs.CL ·
דיאלקט-אמינותיים של דיבור-שפה עם סינתטי פסאודו-דיאלקט אוגמנטציה
Dialect-Robust Speech Language Models with Synthetic Pseudo-Dialect Augmentation
במאמר זה, נציגים שיטה לשיפור יכולת הדיבור-שפה של דגמי ספיק שפה עם סינתטי פסאודו-דיאלקט אוגמנטציה. השיטה כוללת יצירת דיבור-שפה פסאודו-דיאלקטי על ידי תרגום טקסט-לדיבור של LLM-גנריר טקסט-דיאלקטי. השיטה נבדקה על ידי תרגום-דיבור-אנגלי של דיאלקטים שונים, כגון יפני, גרמני וסיני.
תקציר מקורי באנגליתarXiv:2610.09321v1 Announce Type: new Abstract: Speech Language Model (SLM) performance often degrades on dialects due to data scarcity. Conventional text-to-speech (TTS) augmentation struggles to cover diverse dialects as it requires a certain amount of real dialect speech. We propose synthesizing pseudo-dialect speech by converting LLM-generated dialect text via a standard-language TTS model, requiring zero real dialect speech. Additionally, we introduce intermediate standard-text prediction during training, acting as semantic normalization for downstream tasks. We evaluate dialect understanding via dialect-to-English speech translation across Japanese, German, and Chinese dialects. Compared to synthetic standard speech baselines, pseudo-dialect augmentation improves scores for Japanese
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית