כתבה
arXiv cs.CL ·
האם מודלי שפה מבוססי דיבור באמת לומדים מילים?
Do speech foundation models really learn words?
מודלים עצמאיים של דיבור משמשים ביישומים רבים, כולל הכרת דיבור ומודלים של שפה. מחקר זה בודק האם מודלים אלו לומדים מילים באופן עצמאי, ומציע גישה חדשה להבנת הנושא. המחקר מראה כי HuBERT ו-wav2vec 2.0 לומדים ייצוגים של מילים באופן עצמאי.
תקציר מקורי באנגליתarXiv:2609.10434v1 Announce Type: new Abstract: Self-supervised speech foundation models are now used in a wide array of downstream applications, including traditional speech recognition and as the basis for tokens in speech-aware language models. Attempts to understand their usefulness have largely focused on probing their representations' ability to discriminate phonemes and words. However, discriminative ability for words need not imply specialized representation of words per se. Good discrimination of words may be explained by good encoding of word form (phonemes) rather than form-independent word representations encoding identity or syntactic/semantic properties. By partialling out phoneme information using residualization, we show that, in later layers, HuBERT and wav2vec 2.0 do in g
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית