כתבה
arXiv cs.AI ·
מס טוקניזציה: עלות קרוס-לינגואלית של טוקניזציה מילולית עבור שפות הודיות
The Tokenizer Tax: Quantifying and Explaining the Cross-Lingual Cost of Subword Tokenization for Indian Languages
חוקרים בדקו את העלות הקרוס-לינגואלית של טוקניזציה מילולית עבור שפות הודיות. הם מצאו שמודלים כמו GPT-3.5 ו-GPT-4 יוצרים נטל על שפות אלו, אך מודלים רב-לשוניים כמו XLM-R מפחיתים את הנטל.
תקציר מקורי באנגליתarXiv:2607.24276v1 Announce Type: cross Abstract: Large language models (LLMs) process text through subword tokenizers rather than directly reading characters or words. Because these tokenizers are trained predominantly on English-centric corpora, they introduce a systematic and often overlooked disadvantage for many non-English languages. In this work, we quantify this tokenizer tax for Indian languages using the FLORES-200 parallel corpus, measuring tokenization fertility across six widely used tokenizers and fourteen languages. Under cl100k_base (used by GPT-3.5 and GPT-4), Indian languages experience an average 8.0x tokenization tax relative to English, reaching 13.0x for Malayalam, reducing the effective context window to as little as 12% of that available to English users for equival
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית