כתבה
arXiv cs.LG ·
Interleaved Speech Language Models Latently Work In Text
תקציר מקורי באנגליתarXiv:2606.22473v2 Announce Type: replace-cross Abstract: Speech language models (SLMs) increasingly combine speech and text, often by interleaving their tokens within a single sequence. Yet how these two modalities interact in the model's latent space remains unclear. In this work, we analyze interleaved speech--text LMs from different model families and training configurations using three complementary methods. We reveal that these models pass through an implicit latent transcription phase in which the text token matching the spoken word becomes decodable in intermediate layers, despite not being trained for speech recognition. This phenomenon occurs in diverse, natural speech, and intermediate representations also encode likely text continuations. We further show that implicit transcrip
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית