כתבה
arXiv cs.CL ·
שיטה נוירולוגית לפיצול סנדי
A Character-Level Neural Approach to Sinhala Sandhi Splitting
שיטה נוירולוגית לפיצול סנדי: חידוש חשוב ב-NLP הסינהלית
תקציר מקורי באנגליתarXiv:2609.36131v1 Announce Type: new Abstract: Sinhala Sandhi splitting recovers the constituent words or morphemes hidden inside a phonologically merged surface form. The task is important for Sinhala NLP because Sandhi obscures lexical boundaries, but no prior published work has established a neural benchmark for Sinhala Sandhi splitting. We present a character-level sequence-to-sequence study based on SandhiLex, using native Sinhala Unicode input and evaluating recurrent encoder-decoder models for affixational and more complex lexicalized, derivational, and etymological Sandhi. The central challenge is the hard subset lexicalized, derivational, and etymological Sandhi, where our best model, a bidirectional LSTM encoder with a unidirectional LSTM decoder, reaches only 68.40\% exact-matc
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית