Transformer Explainer shows GPT-2 tokenization: 'empowers' splits into two tokens
petrusenko_max · x · 2026-09-22
A look at Transformer Explainer, which runs on GPT-2 small (124M parameters, 50,257-token vocabulary). In its sample prompt, 'Data' and 'visualization' are single tokens, while 'empowers' splits into two — a neat illustration of how irregular tokenization is.
More from Models
- Xiaomi MiMo-v2.6-Pro hands-on: agentic coding score triples, 900tps peak speed — karminski3 · 2026-09-23
- Early user feedback: Claude xhigh/ultracode feels overeager vs plain opus high — MarvinTBaumann · 2026-09-22
- Chart-redrawing capability takes a noticeable leap, months-long recurring test shows — Wattenberger · 2026-09-22
- Musk shows Grok 4.7 doing real engineering at Tesla, shipping FSD features overnight — elonmusk · 2026-09-22
- Dev Praises Qwen3.8-max-preview xhigh, Ran Workloads Nearly 24 Hours on Generous Quota — jasonkneen · 2026-09-22
- Musk Announces Grok 4.7, Cited Demo Builds Interactive 3D Jet Engine in Minutes — elonmusk · 2026-09-22