כתבה
arXiv cs.CL ·
5-Dialects-BN: Unmasking the Impact of Transliteration on Bangla Dialectal LLMs
תקציר מקורי באנגליתarXiv:2609.09964v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved remarkable progress across natural language processing (NLP) tasks, yet their capabilities degrade sharply for low-resource languages and dialectally diverse settings. Bangla, the world's sixth most spoken language, exemplifies this gap: existing resources overwhelmingly target Standard Bangla, leaving its regional dialects without the benchmarks needed to develop or evaluate dialect-aware systems. We address this gap with 5-Dialects-BN, the first multi-annotation Bangla dialect benchmark to align Romanized transliteration with dialectal text, Standard Bangla, English, and subjectivity labels across five regional varieties. The dataset comprises 6,000 manually annotated entries spanning five major di
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית