יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

StreamAlign: Streaming Text-Aligned Speech Tokenization

StreamAlign מאפשר תזמון דיבור-טקסט בזרם, כדי לאפשר יישומים מציאותיים של ייצוג דיבור-טקסט. StreamAlign משלב תזמון דיבור-טקסט באופן אונליין, על ידי חיבור של תזמון RNN-Transducer ברמת התו לתזמון ASR ברמת המילה. StreamAlign נותן תוצאות טובות יותר מאשר תזמונים קיימים, ומציע תזמון דיבור-טקסט בזרם.
תקציר מקורי באנגליתarXiv:2609.09719v1 Announce Type: new Abstract: Text-aligned speech tokenization methods have emerged to better align speech tokens with LLM token spaces, enabling more effective utilization of pretrained LLMs. However, they rely on offline automatic speech recognition (ASR), leading to two key limitations: (i) the need for complete utterances before tokenization, precluding real-time streaming, and (ii) vocabulary mismatch between ASR and LLMs, which reduces acoustic granularity from the subword to the word level. We introduce StreamAlign, a text-aligned speech tokenization framework that enables streaming tokenization for real-time speech-text joint modeling. StreamAlign performs online speech-text alignment by combining character-level RNN-Transducer alignment with word-level ASR guidan
קרא במקור המקורי