Skip to content
important Open-source

Hugging Face has released tokenizers v1 with significantly faster text processing

clearly official source

Hugging Face has released the tokenizers v1 release candidate, which, according to the company, is up to ten times faster than v0.23 while preserving the same API, output and vocabulary. The speedup comes from SIMD optimizations, caching recurring words and native parallelism.

Hugging Face has released tokenizers v1 (as a release candidate) of its tokenization library. According to the company, the new version is up to ten times faster than the previous v0.23, and in some cases the text reports speedups of “even dozens of times”. The output, API, vocabulary and learned merge ranks remain unchanged according to the company, so existing code should work without changes, and the library continues to support the same range of model families as the previous version (BPE, WordPiece and Unigram).

The speedup is achieved through a series of technical changes: splitting the code into separate modules (workspace split), eliminating memory allocations in the main model loop, using SIMD instructions to find split points in text instead of a general-purpose regex engine (a technique called bitcannon), rewriting the merge loop using a linked list in a preallocated buffer, caching recurring words and providing native support for parallel processing across multiple threads without waiting on a single lock.

According to its own account, Hugging Face drew on insights from other open-source tokenization libraries (gigatoken, tiktoken, kitoken, tokie, fastokens, wordchipper, ai-tokenizer) and collaborated with IBM, NVIDIA and the ExecuTorch team on testing across hardware. The company compares the results of the new version with alternatives in a separate repository, tokbench, where the benchmarks can also be run on your own hardware.

The source article is incomplete and does not cover all parts of the original text—further information, including detailed benchmark figures, can be found in the source article.

What changed

Why it matters

Tokenization is often a bottleneck when training on large datasets or serving many requests, which, according to Hugging Face, can leave the GPU idle while it waits for the CPU. Faster tokenization with the API preserved means developers can save compute time without having to change existing code, which is particularly relevant for teams running large-scale ML pipelines.

Two audiences, two different impacts

What this means

01

For individuals

Developers using the tokenizers library in projects with models may, according to Hugging Face, gain significantly faster text encoding and decoding by upgrading to v1 without changing existing code, because the API, output and vocabulary are preserved.

What to do When updating project dependencies, check the tokenizers v1 release candidate and verify compatibility, because Hugging Face states that the API, output and merge ranks are preserved compared with v0.23.
More practical updates →
02

For a business

Companies training on large datasets or serving many concurrent requests may reduce CPU load, which, according to Hugging Face, may previously have limited GPU utilization—but this is an infrastructure saving, not a new product or a direct cost or revenue.

Development
What to decide Check whether internal ML pipelines for training or serving models use the tokenizers library, and consider testing v1 to reduce CPU load when processing large volumes of text.
More business impacts →
benchmark Hugging Face inference infrastruktura tokenizers výkon

Check the original

Event sources

clearly official source · 1 publisher, 0 independent. We count feeds from the same owner only once.

1
Hugging Face Blog primary source · first detected tokenizers v1: encode, decode and scaling, measured