כתבה
arXiv cs.CL ·
צמצום פער הפלט במודלים של שפה מדוברת
Reducing the Output-Mode Gap in Speech Language Models via Joint-Output On-Policy Distillation
חוקרים הצליחו לצמצם את פער הפלט במודלים של שפה מדוברת באמצעות שיטה חדשה. השיטה, הנקראת Joint-Output On-Policy Distillation, משפרת את דיוק התשובות במודלים של שפה מדוברת. החוקרים בדקו את השיטה על מודלים שונים, כולל LLaMA ו-GPT.
תקציר מקורי באנגליתarXiv:2609.15313v1 Announce Type: cross Abstract: Autoregressive generation of interleaved text and acoustic tokens is a common approach to spoken-response generation in speech large language models. Although this design enables streaming generation with explicit textual guidance, generated acoustic tokens become part of the context for subsequent text predictions. Given identical speech inputs, we observe markedly lower answer accuracy for the internal text generated in speech-to-text-and-speech (S2TS) mode than for speech-to-text (S2T) responses. We term this discrepancy the \emph{output-mode gap} (OMG). To reduce OMG, we propose \emph{Joint-Output On-Policy Distillation} (JO-OPD), which distills the model's stronger S2T policy into joint generation using student-generated S2TS trajector
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית