Hugging Face says Tokenizers v1 is up to 30x faster
The release candidate keeps token IDs compatible while rebuilding the hot path around SIMD splitting, reusable buffers and native parallelism.
What changed
Hugging Face published benchmark results for the Tokenizers v1 release candidate. The new path uses a SIMD-oriented splitter, reusable scratch buffers, a thread-local word cache, cheaper BPE merge structures and native multi-threaded encoding.
The project says v1 preserves the same token IDs, vocabulary and merge ranks as v0.23 while changing the implementation underneath.
The headline result
Across ten measured model families, Hugging Face reports 3x to 30x faster encoding on one Apple M4 Max thread and 76% linear scaling across eight workers. The project also published tokbench so developers can rerun the measurements.
Hype check
These are first-party release-candidate benchmarks. Python-call overhead, server CPUs, cache behavior and end-to-end inference pipelines can produce different gains. The implementation is interesting precisely because it is reproducible, so independent results should arrive quickly.