כתבה
arXiv cs.CL ·
סגירת הפער בין דיבור לטקסט עם אודיו מוגבל להתאמה תחומית יעילה
Closing the Speech-Text Gap with Limited Audio for Effective Domain Adaptation in LLM-Based ASR
חוקרים בדקו אם כמויות קטנות של אודיו יכולות לסגור את הפער בין דיבור לטקסט במודלים של LLM-Based ASR. התוצאות הראו שאפילו כמויות קטנות של אודיו משפרות באופן משמעותי את התוצאות. המחקר השתמש במודל LLaMA ובמסגרת LangChain.
תקציר מקורי באנגליתarXiv:2604.06487v3 Announce Type: replace Abstract: Conventional end-to-end automatic speech recognition (ASR) systems rely on paired speech-text data for domain adaptation. Recent LLM-based ASR architectures connect a speech encoder to a large language model via a projection module, enabling adaptation with text-only data. However, this introduces a modality gap, as the LLM is not exposed to the noisy representations produced by the speech projector. We investigate whether small amounts of speech can mitigate this mismatch. We compare three strategies: text-only adaptation, paired speech-text adaptation, and mixed batching (MB), which combines both. Experiments in in-domain and out-of-domain settings show that even limited speech consistently improves performance. Notably, MB using only 1
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית