יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

מעבר לטוקנים אטומיים

Beyond Atomic Tokens: Factorizing Syllables for Language Model Pretraining
פותח טוקניזר פונמי לווייטנאמית וסינית, המגיע לרמת יעילות גבוהה יותר. הטוקניזר מחלק כל הברה לשלושה רכיבים: התחלה, סיום וטון. הדבר מאפשר ייצוג משותף של הברות פונולוגיות קשורות.
תקציר מקורי באנגליתarXiv:2609.21362v2 Announce Type: replace Abstract: Conventional tokenizers represent text as characters or statistically derived subwords, overlooking the internal phonological structure of syllables and often requiring large vocabularies. We introduce \textbf{Phonemic Tokenizer}, a linguistically motivated tokenizer for Vietnamese and Chinese that converts each syllable into IPA and factorizes it into three phonological components: onset, rime, and tone. The three components jointly occupy one contextual position, preserving syllable-level sequence length while enabling representation sharing across phonologically related syllables. Non-phonological and unsupported units are handled through character-level fallback. This deterministic design requires no corpus-dependent vocabulary learni
קרא במקור המקורי