Transformer Explainer shows GPT-2 tokenization: 'empowers' splits into two tokens

petrusenko_max · x · 2026-09-22

A look at Transformer Explainer, which runs on GPT-2 small (124M parameters, 50,257-token vocabulary). In its sample prompt, 'Data' and 'visualization' are single tokens, while 'empowers' splits into two — a neat illustration of how irregular tokenization is.

Original post →

More from Models

Models channel →