AI2's Nature Paper: Byteification Retrofits LLMs to Byte-Level for Under 1% of Pretraining Cost
TheTuringPost · x · 2026-10-09
A new Nature paper from Allen AI introduces "byteification," a two-stage distillation method that retrofits pretrained models (Olmo, Llama 3, Qwen) into byte-level systems using under 1% of the original pretraining budget. By adding 1 byte of lookahead during prefill and predicting fused boundary tokens at decode, the resulting Bolmo, Blama and Bwen models run at practical speeds, crush character-level reasoning, and transfer instruction tuning via task arithmetic.
More from Models
- Test shows Haiku 5.5 triples robot task success with more thinking, GPT-6 Luna stays near zero — ycombinator · 2026-10-10
- Open-EmbeddingGemma thread wraps up: why Kye Gomez reimplements closed-source models — KyeGomezB · 2026-10-10
- Open-EmbeddingGemma: pure PyTorch reimplementation of Google's 740M multimodal embedding model — KyeGomezB · 2026-10-10
- Sources behind the Open-EmbeddingGemma reimplementation: HF model page and Google blog — KyeGomezB · 2026-10-10
- How EmbeddingGemma 2 works: six steps from five modalities to one 768-dim vector — KyeGomezB · 2026-10-10
- Dev claims Gemini 3.8 Flash beats Opus 5.5 for traditional RAG in finance — rickasaurus · 2026-10-10