Meta paper: byte-level models overtake token models at scale, up to 4% ahead
omarsar0 · x · 2026-09-14
- A Meta paper shows byte-level models start behind token models but overtake them as compute grows, demonstrated on distilled 1B models trained on up to 1 trillion bytes.
- The distillation converts a token teacher's logits into byte logits, either approximately (Marginalize-It) or exactly (End-Of-Token).
- Key findings: token models lead at low compute but plateau; byte models reach a higher ceiling, with fitted scaling laws predicting End-Of-Token ends up to 4% ahead of the distilled token model.
- Efficiency gains: byte models match the distilled token model with one-sixth of the training data, and a 256-entry vocabulary cuts teacher-logit storage to about one-fifth.
More from Research
- Open-source GeoGuesser RL environment trains VLMs on visual geolocation with GRPO — HuggingEnvs · 2026-09-15
- Jeff Clune discusses a newly possible, powerful type of RL in MIT Tech Review interview — jeffclune · 2026-09-15
- Jeff Clune on a powerful new type of RL in MIT Tech Review interview — jeffclune · 2026-09-15
- Main approaches for finetuning e2e driving models in close(ish)-loop — abursuc · 2026-09-15
- Atria Dawn Preview: student-heavy team launches research-focused agentic base model — xiaohu · 2026-09-15
- Light Origins' humanoid parkour policy picks walk, vault or climb with onboard sensing only — micoolcho · 2026-09-15