יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

פיגור בין דגמי שפה מבוססי דיבור לדגמי שפה מבוססי טקסט

Quantifying the Generation Modality Gap in Speech-Text Language Models
חוקרים בדקו את הפער בין דגמי שפה מבוססי דיבור לדגמי שפה מבוססי טקסט. הם פיתחו סוויטת הערכה אחידה להשוואת דגמים שונים. התוצאות הראו שדגמי שפה משולבים משפרים את התוכן הסמנטי, אך השיפור אינו אחיד.
תקציר מקורי באנגליתarXiv:2609.14743v1 Announce Type: new Abstract: Pure speech language models often lag behind text and speech-text language models in generating coherent content, but this gap is difficult to quantify because speech and text systems are typically evaluated with different metrics and trained on different data. We study the speech-text modality gap in a family of spoken language models, based on flow matching for continuous acoustic feature generation. We construct a unified generation-based evaluation suite that compares speech-only, text-only, and speech-text language models trained on matched data distributions and evaluated in matched generation settings. We evaluate generated continuations along multiple dimensions: semantic coherence, measured by transcribing generated speech and scorin
קרא במקור המקורי