כתבה
arXiv cs.CL ·
מיצוי שפה-מודע ל-LLM ארבע-לשוני עם פיקוח ASR
Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision
חוקרים פיתחו שיטה חדשה לאימון מודלי LLM רב-לשוניים. השיטה משתמשת במיצוי שפה-מודע ומשפרת ביצועים על נתונים רב-לשוניים. המחקר כולל גם הצגת בנך' Audio-MLQA לבדיקת מודלים.
תקציר מקורי באנגליתarXiv:2603.07025v2 Announce Type: replace Abstract: Speech Large Language Models (LLMs) that understand and follow instructions in many languages are useful for real-world interaction, but are difficult to train with supervised fine-tuning, requiring large, task-specific speech corpora. While recent distillation-based approaches train performant English-only Speech LLMs using only annotated ASR data by aligning text and speech using only a lightweight projector, these models under-perform when scaled to multilingual settings due to language interference in the shared projector. We address this by introducing language-aware distillation using a query bank and a gating network that selects or mixes query tokens using a Q-Former projector. Our approach shows gains of 14% over matched multilin
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית