כתבה
arXiv cs.CL ·
Xiaomi-CocktailASR-1: דו"ח טכני
Xiaomi-CocktailASR-1 Technical Report
Xiaomi-CocktailASR-1 הוא מודל ASR המבוסס על LLM, המאפשר הפרדה וזיהוי דיבור בתוך רעש רקע. המודל משתמש במבנה end-to-end ומסוגל לזהות דיבור יחיד בתוך רעש. הוא גם מסוגל לדחות דגימות שאינן מכילות את הדובר המטרה.
תקציר מקורי באנגליתarXiv:2609.11274v1 Announce Type: cross Abstract: Recently, large language model (LLM) based ASR models have achieved significant progress, yet they generally lack support for multi-speaker scenarios, where the cocktail party problem remains a critical bottleneck for further advancing ASR. Existing TS-ASR methods, including end-to-end architectures with speaker embeddings and latest LLM-based explorations suffer from degraded single-speaker performance and the inability to reject when the target speaker is absent. In this paper, we propose Xiaomi-CocktailASR-1, an LLM-based end-to-end TS-ASR architecture. By utilizing reference speech as voiceprint prompts, it directly transcribes the target speaker's speech without requiring speech separation. Xiaomi-CocktailASR-1 maintains competitive pe
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית