יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

Typhoon ASR: הכרה דיבור תאילנדית עם זמן תגובה נמוך

Typhoon ASR Streaming: Steerable Low-Latency Thai Speech Recognition with Real-Time Shallow Fusion
Typhoon ASR הוא מערכת להכרת דיבור תאילנדית עם זמן תגובה נמוך. המערכת מאפשרת הכוונה של אוצר המילים בזמן פענוח, ללא צורך באימון מחדש. היא משתמשת במודל שפה ובשכבת ערבוב רדודה כדי לשפר את דיוק ההכרה.
תקציר מקורי באנגליתarXiv:2609.14991v1 Announce Type: new Abstract: Open Thai automatic speech recognition (ASR) is dominated by offline, Whisper-based models that read the whole utterance before transcribing, ruling out low-latency uses such as live captioning and voice agents. We present a deployable system for streaming Thai ASR that lets a user steer its vocabulary at decode time, without retraining. A widely used open Thai model, trained with full context, collapses when run as a true stream; we restore streaming with a cache-aware encoder, by converting it or adapting a natively streaming one, and add a shallow-fusion layer that re-ranks candidates inside the streaming decoder with a GPU n-gram language model and phrase boosting. Across two Thai benchmarks and two model sizes, the streaming models stay
קרא במקור המקורי