32 researchers spent 8 months on the most comprehensive survey of tokenization ever
christopher · x · 2026-10-01
A new paper argues that tokenization is a wildly understudied area of language modeling, despite its effects across all of NLP.
Over the past 8 months, 32 tokenizer researchers put together the most comprehensive survey of the field to date, covering how tokenization choices ripple through language models.
More from Models
- 27B at Q5 with full 131k context on one 24GB RTX 3090, 13-17% faster — bjivanovich · 2026-10-01
- Sol 6.1 Day Two Impressions: Strong at Scheduled Tasks and Data Analysis, Cheap — bindureddy · 2026-10-01
- OpenCode's Free Stealth Model Gives Trillions of Tokens Daily, Tested With Blender MCP — sidahuj · 2026-10-01
- Block AttnRes analysis: a quarter of layers dead at depth 32, peak mixing at 20 layers — ziv_ravid · 2026-10-01
- Claude took on the Riemann Hypothesis — here's what actually happened — elsleightholm · 2026-10-01
- GPT-6.1 Sol coming soon to ChatGPT's regular chat mode — mark_k · 2026-10-01