Bolmo recipe: short additional training retrofits models to bytes

allen_ai · x · 2026-10-07

AI2 described its recipe: a short additional training run adapts existing subword models to bytes while keeping the model's core.

New components group bytes into variable-length patches for the model to process, then expand outputs back to byte-level representations to predict the next byte.

Related event: AI2's byte-level language model Bolmo published in Nature with open weights(6 posts)→

Original post →

More from Research

Research channel →