Why byte-level models matter: subword tokenization breaks code and math

allen_ai · x · 2026-10-07

In its Bolmo thread, AI2 explained the motivation: most language models split text into subwords from a fixed vocabulary, obscuring spelling details across writing systems and splitting meaningful units in code or math.

Byte-level models instead work directly with the bytes computers use to represent text.

Related event: AI2's byte-level language model Bolmo published in Nature with open weights(6 posts)→

Original post →

More from Research

Research channel →