כתבה
arXiv cs.CL ·
השפעות שפתיות על תקבול תקן במודלי שפה רב-לשוניים
Language-Specific Effects of Tokenizer Choice in Multilingual Language Models
בחירת תקבול תקן משפיעה על מודלי שפה רב-לשוניים. נמצא שבחירת תקבול תקן חשובה יותר לשפות עם נתוני הכשרה קטנים. נמצא גם שאלוקציה של נתוני הכשרה לשפות עם נתוני הכשרה קטנים, לא תמיד תסייע להן.
תקציר מקורי באנגליתarXiv:2610.12144v1 Announce Type: new Abstract: Tokenizer choice affects multilingual language modeling, but vocabulary capacity is finite and vocabulary size is often constrained: improving representation for some languages often comes at the expense of others. We therefore ask whether tokenizer choice matters equally across languages, a question that the current literature leave unanswered. To this end, we train 123 language models spanning 54 tokenizers. In the main comparison, architecture, training corpus, training-token budget, and optimization are held fixed, so the models differ only in their tokenizer. We find that tokenizer choice matters more for languages with less language-model training data: across the 54 tokenizers, the standard deviation of a language's bits-per-byte (BPB)
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית