יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

Tokka-Bench: ביקורת טוקניזציה ב-100 שפות טבעיות ו-20 שפות תכנות

Tokka-Bench: Evaluating Tokenizers Across 100 Natural and 20 Programming Languages
ביקורת טוקניזציה ב-100 שפות טבעיות ו-20 שפות תכנות. המאמר מציג פלטפורמה חדשה לביקורת טוקניזציה, Tokka-Bench, שמאפשרת ביקורת טוקניזציה ב-100 שפות טבעיות ו-20 שפות תכנות. הפלטפורמה משתמשת בשבעה טוקניזציות BPE (GPT-2, GPT-4, gpt-oss, Llama 3.1, Gemma 3, Qwen3 ו-Kimi K2) ומציגה תוצאות שונות לכל שפה.
תקציר מקורי באנגליתarXiv:2610.08794v1 Announce Type: new Abstract: Large language models rely on subword tokenizers whose quality varies across languages, yet no standardized multi-metric framework exists for broad comparative evaluation. We introduce Tokka-Bench, an open-source framework that evaluates tokenizers on five complementary metrics -- bytes per token, unique token coverage, subword fertility, word-split rate, and vocabulary composition -- across 100 natural languages (30+ scripts) and 20 programming languages, using language-aware segmentation adapted to each writing system. Comparing seven BPE tokenizers (GPT-2, GPT-4, gpt-oss, Llama 3.1, Gemma 3, Qwen3, and Kimi K2) within individual languages, we find that vocabulary allocation strategy matters more than raw vocabulary size, and that programmi
קרא במקור המקורי