CoBPE tokenization cuts sequences 30% vs BPE, better LLMs at same compute
alisawuffles · x · 2026-10-08
A COLM 2026 thread by Yuval Reif shows CoBPE fixes BPE waste: the Hobbit's opening line takes 13 BPE tokens (7 spent on "in", "a", "the" and punctuation) but only 6 with CoBPE, which attaches function words to their words. Across English, sequences are 30% shorter, yielding better LLMs at the same compute.
More from Research
- OpenAI's new result proves 2005 edit-distance embedding optimal; researcher distills proof to 2.5 pages with AI help — thegautamkamath · 2026-10-08
- Experiments show smarter models and higher effort write better LLM-judge evals — danshipper · 2026-10-08
- Tencent's WorkForge scales verifiable training environments for long-horizon work agents — teortaxesTex · 2026-10-08
- Masked Geometric Encoder boosts 3D foundation models via frame dropping and self-distillation — zhenjun_zhao · 2026-10-08
- DensiTok: flow-matching token densification lets frozen feed-forward 3DGS see unseen views — zhenjun_zhao · 2026-10-08
- Warping flat-port views into pinhole perspective for underwater 3D reconstruction — zhenjun_zhao · 2026-10-08