כתבה
arXiv cs.LG ·
קידוד בתים רגיל: רוטינג UTF-8/UTF-16 להפחתת תקרת תקציב תגיות-כרכוב
Universal Byte-Level Encoding: UTF-8/UTF-16 Routing to Reduce Cross-Script Token-Budget Disparities
קידוד חדש מפחית תקרת תקציב תגיות-כרכוב במודלי שפה רב-לשוניים. הצעת קידוד חדשה, UBE, מקידד 1-2-בתים UTF-8 על UTF-8 ו-3-4-בתים UTF-8 דרך UTF-16. זה מוריד את קוצב הקידוד הנמוך ביותר (floor) ל-3-בתים BMP תווים בשפות עם תגיות-כרכוב יקרות (premiums) ללא עלייתו לשפות כבר יעילות בטקסט משולב.
תקציר מקורי באנגליתarXiv:2610.01984v2 Announce Type: replace-cross Abstract: Byte-level byte-pair encoding (BBPE) tokenizers are attractive for multilingual large language models (LLMs) because they cover all Unicode text. In UTF-8-based BBPE, however, many scripts start from a higher fallback cost than English: when no learned merges can be applied, a multibyte character requires multiple byte-derived symbols. We call this worst-case pre-merge cost the encoding floor. A higher floor can increase token counts and per-request cost and shrink usable context. Changing the text encoding can reduce this gap, but a single global encoding can make already-efficient English spans more expensive in mixed-script text. We propose Universal Byte-Level Encoding (UBE), a dual-alphabet tokenizer that keeps 1-2-byte UTF-8 c
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית