AI2's Nature Paper: Byteification Retrofits LLMs to Byte-Level for Under 1% of Pretraining Cost

TheTuringPost · x · 2026-10-09

A new Nature paper from Allen AI introduces "byteification," a two-stage distillation method that retrofits pretrained models (Olmo, Llama 3, Qwen) into byte-level systems using under 1% of the original pretraining budget. By adding 1 byte of lookahead during prefill and predicting fused boundary tokens at decode, the resulting Bolmo, Blama and Bwen models run at practical speeds, crush character-level reasoning, and transfer instruction tuning via task arithmetic.

Original post →

More from Models

Models channel →