32 researchers publish most comprehensive survey of tokenization in NLP
mcmcmcmcmcmcmcmcmc_ · reddit · 2026-10-01
32 tokenizer researchers spent 8 months producing what they call the most comprehensive survey of tokenization, an understudied area affecting all of NLP. It covers algorithms, evaluation, multilinguality, encodings, and theory; potential replacements like latent and visual tokenization; and adjacent topics including constrained generation, token healing, and tokenizer security.
Related event: 32 researchers publish most comprehensive tokenization survey yet(2 posts)→
More from Research
- HCOMP 2026 final session opens with FIBB paper on false information beliefs — windx0303 · 2026-10-01
- Meta-reasoning harness hits 71.5% on ProgramBench, beating Codex by 13.5 points — anirudhg9119 · 2026-10-01
- Rethinking inference-time scaling: what to compute, not just how much — anirudhg9119 · 2026-10-01
- Meta researchers propose Agentic Meta-Reasoning: letting models build their own reasoning graphs — anirudhg9119 · 2026-10-01
- FINGR: a dexterous hand solves a 2×2 Rubik's Cube with continuous finger tricks — HaozhiQ · 2026-10-01
- New lower bounds for kissing numbers: τ₁₉≥12268 and τ₂₁≥30761, open-sourced — felpix_ · 2026-10-01