Byte-level wrapper lets language models read text letter by letter at near-original cost

bravo_abad · x · 2026-10-08

Most LLMs read text as tokens, so the letters inside a token aren't represented separately—explaining why models write fluently but struggle to rearrange letters, a problem that matters in domains like DNA where one letter can change gene function.

Minixhofer and colleagues retrofit an existing pretrained transformer without rebuilding it: small byte-level layers wrap the model, with an encoder reading byte by byte and grouping bytes into patches that initially mimic the original tokens. The central transformer still processes one unit per patch, keeping cost close to the original, and a decoder writes output byte by byte—giving true character-level manipulation ability.

Original post →

More from Research

Research channel →