כתבה
arXiv cs.CL ·
הפעלה מונחה פונמה להכרה עצמית של דיבור
Phoneme-Guided Initialization for LLM-based Speech Recognition
את המאמר זה כתבו חוקרים שפיתחו שיטה חדשה לשיפור יכולת ההכרה העצמית של דיבור במצבי משאבים נמוכים. השיטה, הנקראת ‘הפעלה מונחה פונמה’, משתמשת בפונמות (צורות יסוד של צלילי השפה) כדי לשפר את יכולת המודל להבחין במילים.
תקציר מקורי באנגליתarXiv:2610.08994v1 Announce Type: cross Abstract: Speech large language models (speech LLMs) perform well on automatic speech recognition (ASR) when sufficient paired speech-text data is available, but their performance degrades in low-resource settings. A cascaded pipeline that performs speech-to-phoneme (S2P) conversion followed by phoneme-to-grapheme (P2G) conversion has been shown to outperform end-to-end speech LLMs in this regime, suggesting that phoneme-mediated processing is beneficial when paired data is scarce. We propose \textit{phoneme-guided initialization}, a simple method that uses this insight within an end-to-end framework: we pre-train the audio encoder on S2P and the LLM on P2G tasks, then connect them and fine-tune the full model end-to-end on the target ASR task. Exper
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית