כתבה
arXiv cs.CL ·
הפחתת הבלבול בין ניב-לשון בייצוגי דיבור-עצמי לזיהוי שפה
Mitigating Accent-Language Confusion in Self-Supervised Speech Representations for Language Identification
במאמר זה, המחברים מציגים פתרון לבעיה של הבלבול בין ניב-לשון בייצוגי דיבור-עצמי לזיהוי שפה. הם מציגים פרויקציה גאומטרית שמורידה את הטיה-לשון ה-L1 מהייצוגי הדיבור-עצמי, ומשפרת את זיהוי השפה לדיבור-עצמי של L2.
תקציר מקורי באנגליתarXiv:2610.09486v1 Announce Type: cross Abstract: Spoken language identification (LID) aims to recognize the target language regardless of accent. In practice, however, LID models fine-tuned from self-supervised speech representations frequently confuse accents with languages, misclassifying non-native (L2) speech as the speaker's first language (L1). We show that non-native speech representations lie between native target-language and native L1 poles, causing systematic misclassification. To address this, we introduce a geometric projection that estimates an L1-bias direction solely from native speech and removes it before the frozen LID head. Across five MMS-LID models and non-native corpora, this projection substantially improves target language identification for L2-accented speech whi
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית