כתבה
arXiv cs.AI ·
The Price of Token Boundaries: Compression Certificates and Prediction
תקציר מקורי באנגליתarXiv:2609.35869v1 Announce Type: new Abstract: Pre-tokenisation restricts which text fragments can become prediction units, but its compression cost is obscured when tokenisers are compared only under the same boundaries. We measure this cost by bounding the minimum token count from both sides, with and without a regular-expression boundary rule. Nonnegative prices on token occurrences yield a lower bound through shortest paths and vocabulary-budget selection; maximising over all prices recovers the linear programming relaxation, and an independent integer checker certifies the reported values. On English Wikipedia, boundaries increase the optimal token count by 28.3--36.8\%. Byte pair encoding lies 2.1\% above the constrained lower bound, but 10.9\% above the unrestricted bound. Compress
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית