כתבה
arXiv cs.CL ·
Nuha-Speech: בניית מודלים כלליים לשפה ערבית
Nuha-Speech: Building General-Purpose Arabic Speech-LLMs
Nuha-Speech היא יוזמה לפיתוח מודלים כלליים לשפה ערבית. היא כוללת בניית מאגר נתונים גדול ואימון מודלים Qwen-Omni. המטרה היא לספק תשתית למודלים ערביים.
תקציר מקורי באנגליתarXiv:2609.11892v1 Announce Type: new Abstract: As Speech Large Language Models (speech-LLMs) become increasingly multilingual, Arabic remains significantly underrepresented, highlighting the need for dedicated infrastructure to train and evaluate Arabic speech-LLMs. To address this gap, we introduce Nuha-Speech, a comprehensive initiative to develop general-purpose Arabic speech-LLMs spanning dataset construction, model training, and systematic evaluation. Specifically, we constructed a large-scale Arabic Speech Question-Answering (SQA) corpus comprising over 1.5 million training samples to allow instruction tuning over a broad range of core speech tasks. Then, the corpus was used for supervised fine-tuning based on Qwen-Omni model variants at different scales. Finally, we designed an eva
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית