Hugging Face Tokenizers
Hugging Face's Rust tokenizer library for research and production inference.
What happened lately
Hugging Face Tokenizers timeline
What it is
Tokenizers is Hugging Face’s Rust library for converting text into the token IDs consumed by language models. It supports production and research workloads across multiple tokenizer families.
What v1 changes
The v1 release candidate rewrites major parts of the encoding path with reusable buffers, SIMD-oriented splitting, a word cache, native parallelism and a cheaper merge loop. Hugging Face says output IDs remain compatible with v0.23.
The measured claim
Hugging Face reports encoding improvements of 3x to 30x across ten model families on one Apple M4 Max thread, with 76% linear scaling across eight workers. The accompanying tokbench repository is intended to make those results reproducible.
Watch point
The numbers come from the project’s own benchmark and a release candidate. Independent results on server CPUs, Python bindings and real inference pipelines will show how much of the gain survives outside the Rust benchmark harness.