כתבה
arXiv cs.CL ·
קידוד בתים רגיל: רוטינג UTF-8/UTF-16 להפחתת תקרת תקציב טוקן-צלע בין-סקריפט
Universal Byte-Level Encoding: UTF-8/UTF-16 Routing to Reduce Cross-Script Token-Budget Disparities
קידוד בתים רגיל למודלי LLM רב-לשוניים מפחית תקרת תקציב טוקן-צלע בין-סקריפט
תקציר מקורי באנגליתarXiv:2610.01984v2 Announce Type: replace Abstract: Byte-level byte-pair encoding (BBPE) tokenizers are attractive for multilingual large language models (LLMs) because they cover all Unicode text. In UTF-8-based BBPE, however, many scripts start from a higher fallback cost than English: when no learned merges can be applied, a multibyte character requires multiple byte-derived symbols. We call this worst-case pre-merge cost the encoding floor. A higher floor can increase token counts and per-request cost and shrink usable context. Changing the text encoding can reduce this gap, but a single global encoding can make already-efficient English spans more expensive in mixed-script text. We propose Universal Byte-Level Encoding (UBE), a dual-alphabet tokenizer that keeps 1-2-byte UTF-8 charact
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית