כתבה
arXiv cs.CL ·
Indic DiarBench: בנך' מרובת לשונות לדיאריזציה ו-ASR
Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages
Indic DiarBench הוא בנך' לדיאריזציה ו-ASR רב-לשוני ל-22 שפות הודיות. הוא כולל 108 שעות אודיו רב-דוברים. המחקר בודק מודלים מובילים, כולל API לדיבור ומודלים רב-מודאליים.
תקציר מקורי באנגליתarXiv:2607.23808v1 Announce Type: new Abstract: In this work, we introduce Indic DiarBench, a speaker diarization and ASR benchmark dataset spanning all 22 scheduled languages of India. This corpus comprises approximately 108 hours of natural multi-speaker audio from near-field meetings, far-field recordings, and in-the-wild audios. All annotations are human-corrected with time-aligned speaker attributed transcriptions. The dataset captures conversational nuance prevalent in Indian speech, such as English code-mixing, dialectal variation, and frequent speaker overlap. To establish a baseline for joint ASR and diarization capabilities we evaluate leading systems including commercial speech APIs and multimodal large language models. Indic DiarBench is released as an open-access resource to a
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית