יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

שיטה נוירולוגית לפיצול סנדי

A Character-Level Neural Approach to Sinhala Sandhi Splitting
שיטה נוירולוגית לפיצול סנדי: חידוש חשוב ב-NLP הסינהלית
תקציר מקורי באנגליתarXiv:2609.36131v1 Announce Type: new Abstract: Sinhala Sandhi splitting recovers the constituent words or morphemes hidden inside a phonologically merged surface form. The task is important for Sinhala NLP because Sandhi obscures lexical boundaries, but no prior published work has established a neural benchmark for Sinhala Sandhi splitting. We present a character-level sequence-to-sequence study based on SandhiLex, using native Sinhala Unicode input and evaluating recurrent encoder-decoder models for affixational and more complex lexicalized, derivational, and etymological Sandhi. The central challenge is the hard subset lexicalized, derivational, and etymological Sandhi, where our best model, a bidirectional LSTM encoder with a unidirectional LSTM decoder, reaches only 68.40\% exact-matc
קרא במקור המקורי