כתבה
arXiv cs.AI ·
Data Scale, Not Latency, Shapes Cross-Lingual Encoder Transfer in Streaming ASR
קנה מידה הנתונים, ולא עיכוב, משפיע על תרגום מקביל-שפה של מודלי ASR בזרימה.
תקציר מקורי באנגליתarXiv:2606.24169v2 Announce Type: replace Abstract: Adapting a streaming speech recognition model to a new language requires choosing between two plausible warm starts: a multilingual (ML) encoder or an English-only (EN) encoder. The common intuition is that the multilingual encoder should help most at low data, but it is unclear how long that advantage persists, whether tight streaming latency amplifies it, and whether it survives deployment quantization. We answer these questions with a controlled sweep of a 0.6 B-parameter cache-aware FastConformer transducer across eight European languages, up to five target-language data scales (100 h to 2500 h), three streaming tiers plus offline decoding, and up to four public test sets. The main result is that multilingual initialization is a data-
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית