Tokenization isn't why LLMs miscount letters: models spell characters with 100% accuracy

maksym_andr · x · 2026-09-21

Challenging the popular explanation that tokenization causes frontier LLMs to miscount letters, the author notes that while models only see token IDs during training, any frontier model can spell text character-by-character with 100% accuracy—perfectly mapping token IDs to characters. The bottleneck is likely the counting/aggregation operation itself, not blindness to exact characters; tokenization has weird effects but doesn't make LLMs character-blind.

Original post →

More from Models

Models channel →