Tokenizers v1 heads to SIMD refactors after claims of 500–1000x speedups

vanstriendaniel · x · 2026-07-22

Tokenizer refactor moves to SIMD, with v1 coming soon

The post says the team has been refactoring tokenizers to use SIMD in pre-tokenization and normalization, while redesigning the core structure of BPE and WordPiece for speed. A Tokenizers v1 release is on the way.

The quoted thread also points to Gigatoken, which claims to be roughly 500–1000x faster than HuggingFace tokenizers and about 100x faster than OpenAI’s tiktoken for most tokenizer definitions on most machines.

Original post →

More from Infra

Infra channel →