Why tokenizers still resist end-to-end optimization despite years of pretraining

seanmcdonaldxyz · x · 2026-07-29

Paul Cal argues that text token representation has not been heavily optimized because the feedback loop is too slow and expensive: frontier pretraining is rare, and bad tokenization choices may only show up at scale.

He also notes that tokenizers are fundamentally awkward to optimize end-to-end:

His conclusion: there is still no established, frontier-viable end-to-end tokenizer optimizer that can simultaneously minimize tokens and preserve performance across tasks.

Related event: High Pre-training Costs Hinder Tokenizer Optimization(2 posts)→

Original post →

More from Research

Research channel →