כתבה
arXiv cs.CL ·
Neyshekar: An Open Persian Read-Speech Corpus for Automatic Speech Recognition
תקציר מקורי באנגליתarXiv:2609.14542v1 Announce Type: new Abstract: Neyshekar is presented as an open Persian read-speech corpus designed for coverage of both formal and informal language, named entities, and longer utterances. In version 6, 62,279 validated recordings totalling 99.02 hours are provided from 190 contributors, with 34,541 distinct recorded prompts. The prompt pool was assembled from human-written material, contextualised homographs, and reviewed language-model-generated text. Text entries were normalised with the shekar library, which supports both formal and informal Persian, and every submitted recording was reviewed against a common validation rubric. About 24% of released clips are classified as informal by an automatic classifier; these register labels are not human-validated. Item-level
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית