יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

Qwen-Audio-3.0-ASR: תיעוד טכני

Qwen-Audio-3.0-ASR Technical Report
Qwen-Audio-3.0-ASR הוא מערכת ASR מבוססת LLM של Qwen שתוכננה לקדם דרישות ייצור, כולל הכרה של 30 שפות ו-16 גרסאות דיאלקטיות של סינית. המערכת תומכת בהכרה של ישויות תעשייתיות, הגדרת קלטים חיצוניים, ומודלי תצוגה יחיד-פס.
תקציר מקורי באנגליתarXiv:2609.07549v2 Announce Type: replace Abstract: In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model scaling, and deep integration with large language models (LLMs). However, bridging the gap between academic benchmark performance and real-world production utility remains a persistent challenge, particularly in handling diverse regional dialects, dynamic entities and hotwords, long-range contextual information, and disfluent spontaneous speech. In this report, we present Qwen-Audio-3.0-ASR, a Mixture-of-Experts (MoE) LLM-based ASR system designed to address these production demands through a unified, instruction-following framework. The model is built upon the Qwen backbone, and is tra
קרא במקור המקורי