יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

BaltiVoice: מאגר שיחה ומערכת ASR מותאמת

BaltiVoice: A Speech Corpus and Fine-tuned Whisper ASR System for the Balti Language
BaltiVoice הוא מאגר שיחה בשפה הבלטית, עם 16.8 שעות של הקלטות. המחברים פיתחו מערכת ASR מותאמת באמצעות דגם Whisper של OpenAI, עם שגיאת מילים של 24.78%.
תקציר מקורי באנגליתarXiv:2606.03504v3 Announce Type: replace Abstract: We present BaltiVoice, a 16.8-hour read-speech corpus for Balti (ISO 639-3: bft), a Tibetic language spoken in Gilgit-Baltistan, Pakistan, with no prior publicly available ASR resources. The corpus contains 10,060 validated utterances in native Nastaliq script, derived from Mozilla Common Voice recordings. Fine-tuning OpenAI Whisper-small yields a Word Error Rate (WER) of 24.78% and a Character Error Rate (CER) of 8.30% after training for 5 epochs (3,000 steps) on the 538-utterance speaker-disjoint validation set, down from a zero-shot baseline of 159.19% WER and 152.52% CER. A Whisper-base fine-tuned on the same data achieves 44.54% WER and 15.61% CER, confirming that model capacity matters for this low-resource setting. The dataset, fin
קרא במקור המקורי