יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

פחות מילים, לא פחות סמנים: מדידת עלויות תקיפה סנסקריטיות לפי רעיון

Fewer Words, Not Fewer Tokens: Measuring the Sanskrit Tokenization Penalty per Proposition
במאמר זה נבחן את העלויות של תקיפה סנסקריטית בהשוואה לאנגלית, ובפרט את השפעת גודל הווקאבולרי על התוצאות.
תקציר מקורי באנגליתarXiv:2609.12960v1 Announce Type: new Abstract: Sanskrit fuses case, number, person and tense into word endings and chains clauses into compounds, so it is information-dense per word. Whether that density survives subword tokenization is a separate question, to be asked per unit of meaning rather than per word. On identical FLORES-200 devtest content, Sanskrit costs 1.774-2.187 times the English tokens under deployed tokenizers with vocabularies of 200,019 ids or more, but only 1.325-1.353 times the Hindi tokens. Against a deployed English tokenizer, Sanskrit-trained BPE arms then look cheaper per proposition than English on contemporary prose (0.887). Against a matched English control, the same algorithm and vocabulary trained on the English side of the same corpus, that flip disappears:
קרא במקור המקורי