Tokenization isn't why LLMs miscount letters: models spell characters with 100% accuracy
maksym_andr · x · 2026-09-21
Challenging the popular explanation that tokenization causes frontier LLMs to miscount letters, the author notes that while models only see token IDs during training, any frontier model can spell text character-by-character with 100% accuracy—perfectly mapping token IDs to characters. The bottleneck is likely the counting/aggregation operation itself, not blindness to exact characters; tokenization has weird effects but doesn't make LLMs character-blind.
More from Models
- Report: DeepSeek training a 2T-param model with 8T planned, Huawei chips due late 2026 — Hesamation · 2026-09-21
- Codex PRO+ Users Report Usage Draining Much Faster Since Weekly Reset — Next_Technology6361 · 2026-09-21
- Open-source finetune project laya hits 6.9k stars with free Kaggle 2xT4 notebook — ojasvi_yadav · 2026-09-21
- jev-reranker edges out ruri-v3 on four of five Japanese retrieval benchmarks — amaarora · 2026-09-21
- New model's vuln stats look same as before: many found, few exploited in the wild — xeophon · 2026-09-21
- Claude now makes typos in Russian, continuing output quality complaints — StewartalsopIII · 2026-09-21