כתבה
arXiv cs.CL ·
LMSpell: Spell Correction with Pre-Trained Language Models
תקציר מקורי באנגליתarXiv:2512.05414v4 Announce Type: replace Abstract: Spell correction is still a challenging problem for many languages, especially low-resource languages (LRLs). While pre-trained language models (PLMs) have been employed for spell correction, there has been no proper comparison across PLMs. We present the first empirical study on the effectiveness of the three types of PLMs for spell correction across multiple languages, including low-resource languages. We show that even relatively small PLMs such as the 270M-parameter Gemma 3 and mBART50, when fine-tuned on a dataset of only 5k sentences, can outperform rule-based spell correctors, highlighting a practical pathway for building effective spell correction systems with limited data. We also present a case study with Sinhala to shed light o
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית