יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

קידוד ברמת בייט אוניברסלי

Universal Byte-Level Encoding: UTF-8/UTF-16 Routing to Reduce Cross-Script Token-Budget Disparities
קידוד ברמת בייט אוניברסלי (UBE) הוא פתרון חדש שמפחית את הפערים בתקציב טוקנים בין שפות. UBE משלב בין קידוד UTF-8 ו-UTF-16 כדי להפחית את עלות הטוקנים בשפות עם תקציב טוקנים גבוה. הפתרון הזה יכול לשפר את ביצועי המודלים הרב-לשוניים.
תקציר מקורי באנגליתarXiv:2610.01984v1 Announce Type: cross Abstract: Byte-level byte-pair encoding (BBPE) tokenizers are attractive for multilingual large language models (LLMs) because they cover all Unicode text. In UTF-8-based BBPE, however, many scripts start from a higher fallback cost than English: when no learned merges can be applied, a multibyte character requires multiple byte-derived symbols. We call this worst-case pre-merge cost the encoding floor. A higher floor can increase token counts and per-request cost and shrink usable context. Changing the text encoding can reduce this gap, but a single global encoding can make already-efficient English spans more expensive in mixed-script text. We propose Universal Byte-Level Encoding (UBE), a dual-alphabet tokenizer that keeps 1-2-byte UTF-8 character
קרא במקור המקורי