יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

Morpheus: מפריד מורפולוגיה עצמאי לטורקית

Morpheus: A Morphology-Aware Neural Tokenizer and Word Embedder for Turkish
Morpheus הוא מודל נוירוני המפריד מורפולוגיה עבור השפה הטורקית. הוא משמש הן כמפריד מילים והן כמייצר וקטורים של מילים. Morpheus משיג את התוצאות הטובות ביותר בקטגוריות מסוימות, כגון זיהוי מורפולוגיה ואחזור לקסיקלי.
תקציר מקורי באנגליתarXiv:2606.18717v2 Announce Type: replace-cross Abstract: Turkish is agglutinative: meaning is carried by morphemes, yet the subword tokenizers that drive modern language models split words by corpus statistics, fragmenting semantically loaded suffixes and -- in the case of WordPiece and rule-based analyzers -- failing to decode their output back to the original text. This paper presents \textbf{Morpheus}, a neural morpheme-boundary model for Turkish that is at once a lossless, morphology-aware tokenizer and a word-embedding producer. A differentiable Poisson-binomial dynamic program turns per-character boundary probabilities into soft morpheme memberships during training and exact segments at inference, with no string normalization, so $\mathrm{decode}(\mathrm{encode}(w)) = w$ holds by co
קרא במקור המקורי