יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

מס נסתר של שפה: תכופות כסף של צרפתית ושפות אזוריות ב-2026 LLM Tokenizers, ומודל צרפתי-משופר

The Invisible Language Tax: Token Premiums of French and Regional Languages in 2026 LLM Tokenizers, and a French-Optimized Prototype
צרפתית ושפות אזוריות שולמות יותר ב-LLM Tokenizers עקב תכופות כסף. המחברים חקרו זאת ופיתחו מודל צרפתי-משופר. הם גילו כי צרפתית דורשת 31% ל-58% יותר תגים מאנגלית, ושפות אזוריות של צרפת דורשות כ-1.6 ל-3.3 פעמים יותר מאנגלית. הם גם פיתחו מודל צרפתי-משופר, Baracoda FR v1.2, שמשתמש ב-11.5% פחות תגים מ-Tekken על צרפתית ו-3.7% פחות על אנגלית.
תקציר מקורי באנגליתarXiv:2609.39001v1 Announce Type: new Abstract: LLM services are billed per token and context windows are measured in tokens, yet the number of tokens needed for the same content varies across languages. We measure this token premium on seven tokenizers of widely used 2026 models (OpenAI o200k, Llama 3, Qwen3, DeepSeek V3/V4, Gemma 3, Mistral Tekken, and the Claude generation-5 tokenizer via Anthropic's counting API) on NTREX-128 (124 non-English reference translations) and on the Universal Declaration of Human Rights for regional languages. French requires 31% to 58% more tokens than English, whereas Simplified Chinese ranges from 5% fewer to 40% more and is cheaper than French on six of the seven tokenizers. Regional and overseas languages of France pay roughly 1.6 to 3.3 times the Engli
קרא במקור המקורי