כתבה
arXiv cs.CL ·
מפה תקינה למילים קוראניות: Uthmani ל-Standard
A Corpus-Aligned Uthmani-to-Standard Quranic Word Mapping and a Deterministic Recitation Validator
מפה תקינה למילים קוראניות: Uthmani ל-Standard. פיתוח ואימות של תקינה קוראנית.
תקציר מקורי באנגליתarXiv:2609.14967v1 Announce Type: new Abstract: Quranic text is distributed in two orthographic forms that are byte-level distinct: the Uthmani script used in every printed mushaf, and the Standard (Imla'i) Arabic form that every mainstream Arabic NLP tool is built for. The gap is concentrated in one Unicode character, U+0670 (superscript alef), which appears in some of the most frequently recited words in the Quran and is silently mishandled by general-purpose Arabic normalizers. We release a 2,290-pair, corpus-aligned Uthmani-to-Standard word mapping constructed by aligning the complete 6,236-verse Quran across both orthographic forms, together with a seven-step text normalization pipeline built on it. Normalizing both forms of all 6,236 verses through that pipeline yields identical stri
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית