CoBPE tokenization cuts sequences 30% vs BPE, better LLMs at same compute

alisawuffles · x · 2026-10-08

A COLM 2026 thread by Yuval Reif shows CoBPE fixes BPE waste: the Hobbit's opening line takes 13 BPE tokens (7 spent on "in", "a", "the" and punctuation) but only 6 with CoBPE, which attaches function words to their words. Across English, sequences are 30% shorter, yielding better LLMs at the same compute.

Original post →

More from Research

Research channel →