יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה MarkTechPost ·

גיגאטוקן: מימוש Rust לטוקניזציה ב-24.53 GB/s

Meet Gigatoken: A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s, up to 989x Faster than HuggingFace Tokenizers
גיגאטוקן הוא יישום טוקניזציה ב-Rust שמהיר פי 989 מטוקניזטורים של HuggingFace. הוא תומך ב-23 משפחות טוקניזטורים, כולל GPT-2 ו-Llama. גיגאטוקן משיג ביצועים של 24.53 GB/s.
תקציר מקורי באנגליתTokenization is the one part of the language modeling stack that almost nobody profiles. Gigatoken , released by Marcel Rød (a PhD student from Stanford) under an MIT license, argues that this was a mistake. The library encodes text at gigabytes per second on a single machine, against baselines that are already multithreaded Rust. The GPT-2 tokenizer benchmarking yields remarkable results: evaluated on the 11.9 GB owt_train.txt corpus using a 144-core AMD EPYC 9565 dual-socket setup, Gigatoken processes data at a staggering 24.53 GB/s . In comparison, OpenAI’s tiktoken achieves 36.0 MB/s, while HuggingFace tokenizers registers at 24.8 MB/s on the identical hardware configuration. These marks demonstrate performance advantages of 681x and 989x, respectively. On an Apple M4 Max with 16
קרא במקור המקורי