Tokenization: A Survey for Modern NLP lands on alphaXiv, mapping subword trade-offs
yuvalpi · x · 2026-10-01
A large collaborative survey, 'Tokenization: A Survey for Modern NLP' by Marco Cognetta, Yuval Pinter and dozens of co-authors, is out and trending at #12 on alphaXiv.
- It covers tokenization — the hidden first step of every LLM pipeline — focusing on subword tokenization: it solves OOV vs word-level models while bounding vocab and embedding size, and shortens sequences vs character-level models.
- It also catalogs the flaws: rigid token-building schemes that don't generalize across languages, segmentation ambiguities enabling adversarial attacks, and the need to explicitly balance token allocation between high- and low-resource languages.
More from Research
- Starship Reaches Orbit: How Massive Payloads Could Accelerate Space Telescopes by Decades — misovalko · 2026-10-01
- Google AI proposes RRSI to stop recursive self-improving agents from overfitting benchmarks — burkov · 2026-10-01
- Reddit poster touts 'IQRAX' LLM architecture layer claiming 11.5x lower cost, 48.3x faster — no code released — ThirdCultureMisfit · 2026-10-01
- DeepMind publishes SynthID Bio in Nature, watermarks AI-designed proteins and open-sources the tools — demishassabis · 2026-10-01
- Iso Shows Drug-Design Agent Visualized via Attention Heatmap Over Chemical Space — CatAstro_Piyush · 2026-10-01
- Evaluating One Model on ProgramBench Now Costs Over $10k — jyangballin · 2026-10-01