כתבה
arXiv cs.CL ·
מדידת הפער בין המודלים של דיבור-טקסט
Quantifying the Generation Modality Gap in Speech-Text Language Models
במאמר זה נחקר הפער בין המודלים של דיבור-טקסט בתחום השפה. המחברים חקרו את המודלים של דיבור-טקסט ומצאו שהם משפרים את התכונות הסמנטיות של הדיבור.
תקציר מקורי באנגליתarXiv:2609.14743v2 Announce Type: replace Abstract: Pure speech language models often lag behind text and speech-text language models in generating coherent content, but this gap is difficult to quantify because speech and text systems are typically evaluated with different metrics and trained on different data. We study the speech-text modality gap in a family of spoken language models, based on flow matching for continuous acoustic feature generation. We construct a unified generation-based evaluation suite that compares speech-only, text-only, and speech-text language models trained on matched data distributions and evaluated in matched generation settings. We evaluate generated continuations along multiple dimensions: semantic coherence, measured by transcribing generated speech and sc
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית