כתבה
arXiv cs.CL ·
TutlAit v1: קורפוס דיבור מורוקני-טמאז'ית עם תרגום ערבי ותווי דיבור מקומיים
TutlAit v1: a crowdsourced Moroccan Tamazight speech dataset with Arabic transcriptions and regional accent labels
קורפוס דיבור חדש של טמאז'ית מורוקנית עם תרגום ערבי ותווי דיבור מקומיים. הקורפוס כולל 13,384 קטעי דיבור וניתן לשימוש בתחומים שונים, כגון הכרה של דיבור, תרגום שפות והזיהוי של תווי דיבור.
תקציר מקורי באנגליתarXiv:2609.38219v1 Announce Type: new Abstract: Tamazight (Amazigh) is, together with Arabic, one of the two official languages of Morocco, yet it remains severely under-resourced for speech technology: pub licly available labelled audio is scarce, generally lacks information on the regional variety spoken, and is often of uneven transcription quality. This article describes the TutlAit dataset, a corpus of Moroccan Tamazight speech paired with Modern Standard Arabic text and explicit regional accent labels. The data were collected with TutlAit, a purpose-built crowdsourcing web application (React 18 front end, Django 5 / Django REST Framework back-end, PostgreSQL database). Native speakers recruited through targeted LinkedIn and Instagram campaigns created an account, declared their regio
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית