כתבה
arXiv cs.CL ·
מחיר גבולות טוקנים
The Price of Token Boundaries: Compression Certificates and Prediction
חוקרים בדקו את העלות של גבולות טוקנים על ביצועי דחיסה וניבוי. הם מצאו שגבולות מגדילים את מספר הטוקנים האופטימלי ב-28.3-36.8%. המחקר בודק את השפעת גבולות על דחיסה וניבוי ב-12 שפות.
תקציר מקורי באנגליתarXiv:2609.35869v1 Announce Type: cross Abstract: Pre-tokenisation restricts which text fragments can become prediction units, but its compression cost is obscured when tokenisers are compared only under the same boundaries. We measure this cost by bounding the minimum token count from both sides, with and without a regular-expression boundary rule. Nonnegative prices on token occurrences yield a lower bound through shortest paths and vocabulary-budget selection; maximising over all prices recovers the linear programming relaxation, and an independent integer checker certifies the reported values. On English Wikipedia, boundaries increase the optimal token count by 28.3--36.8\%. Byte pair encoding lies 2.1\% above the constrained lower bound, but 10.9\% above the unrestricted bound. Compre
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית