כתבה
arXiv cs.CL ·
Tokka-Bench: ביקורת טוקניזציה ב-100 שפות טבעיות ו-20 שפות תכנות
Tokka-Bench: Evaluating Tokenizers Across 100 Natural and 20 Programming Languages
ביקורת טוקניזציה ב-100 שפות טבעיות ו-20 שפות תכנות. המאמר מציג פלטפורמה חדשה לביקורת טוקניזציה, Tokka-Bench, שמאפשרת ביקורת טוקניזציה ב-100 שפות טבעיות ו-20 שפות תכנות. הפלטפורמה משתמשת בשבעה טוקניזציות BPE (GPT-2, GPT-4, gpt-oss, Llama 3.1, Gemma 3, Qwen3 ו-Kimi K2) ומציגה תוצאות שונות לכל שפה.
תקציר מקורי באנגליתarXiv:2610.08794v1 Announce Type: new Abstract: Large language models rely on subword tokenizers whose quality varies across languages, yet no standardized multi-metric framework exists for broad comparative evaluation. We introduce Tokka-Bench, an open-source framework that evaluates tokenizers on five complementary metrics -- bytes per token, unique token coverage, subword fertility, word-split rate, and vocabulary composition -- across 100 natural languages (30+ scripts) and 20 programming languages, using language-aware segmentation adapted to each writing system. Comparing seven BPE tokenizers (GPT-2, GPT-4, gpt-oss, Llama 3.1, Gemma 3, Qwen3, and Kimi K2) within individual languages, we find that vocabulary allocation strategy matters more than raw vocabulary size, and that programmi
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית