Hugging Face ships tokenizers v1, often tens of times faster than v0.23
ariG23498 · x · 2026-09-30
Hugging Face released tokenizers v1 with a detailed engineering writeup. As models get faster, tokenization can starve GPUs — massive training runs, concurrent serving, and long-input workloads expose it as a bottleneck.
Key points:
- v1 is often tens of times faster than v0.23, with encode/decode and scaling measured
- Patches and help from IBM, NVIDIA, and the ExecuTorch team
- Refactor borrows proven ideas from open-source ecosystem: gigatoken, tiktoken, kitoken, and others
- Team signals the library is now worth contributing to
The article also frames tokenization as the natural starting point for learning about LLMs.
More from Infra
- Quantized softmax attention pretraining: only +0.004 nats loss gap at K=16 with the right calibration — illinois · 2026-09-30
- xLLM training infra open-sourced with xattn attention backend and xBridges toolkit — HongyiWang10 · 2026-09-30
- Auto-research loop on 120 B300s finds 40% Kimi K3 inference gain for $9,176 — bookwormengr · 2026-09-30
- Cerebras to bring 'world's fastest inference' to General Compute — beffjezos · 2026-09-30
- 8% of Asia-to-US air freight is now data center parts — 30 full freighters a day — yacineMTB · 2026-09-30
- mradermacher quants get Gemma 26B to 75 tok/s on 2x RTX 4060 8GB — Spiritual_Impress_30 · 2026-09-30