toks runs its tokenizer in pure assembly: 1.8 GB/s on 8 Zen 5 cores
cephaloform · x · 2026-10-07
actualinc's toks implements most of its computation in pure assembly, encoding a 16 MiB input at 1.6 GB/s on 8 GB10 cores and 1.8 GB/s on 8 Zen 5 cores — with exactly the ids of one serial call, i.e. parallelized without changing tokenization output.
More from Infra
- Free Online Conference All Day AI Set for Oct 22 With Talks on SLMs and Local Inference — FikoFox · 2026-10-07
- Best $4,000 Local LLM Rig? Weighing R9700s, Strix Halo, and 6x Arc B60 — BinaryGrind · 2026-10-07
- Paris' Tour Montparnasse Plans 120,000 Nvidia B300s Delivering 540 EFLOPS at 304MW — IgorCarron · 2026-10-07
- Qwen 27B at ~18 tok/s on Just 12GB VRAM + 8GB RAM: Full Recipe Released — bodhi371 · 2026-10-07
- 20+ Hand-Tuned ROCm Kernels Nearly Double Qwen 27B Throughput on 4x 7900XTX — NoFee9147 · 2026-10-07
- PyTorch Replaces CUDA with FBTriton for Embedding Kernels: 1.28x Faster Forward, 2x Backward — PyTorch · 2026-10-07