arXiv: distilling byte models breaks the token ceiling, up to 4% better asymptotically
alex_verem · x · 2026-10-10
The paper "Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models" (Meta FAIR + UW, with Luke Zettlemoyer) presents the first large-scale study of byte-level language models.
- Two methods to convert token logits to byte logits: approximate Marginalize-It and exact End-Of-Token (EOT)
- Systematic sweep of 1B-parameter decoder-only models trained on up to 1 trillion bytes, varying tokenization scheme and objective (distillation vs cross-entropy)
- Across eight benchmarks (QA, generation, translation): token models lead in the low-FLOP regime but plateau; byte models start worse yet surpass them with more compute, reaching a higher ceiling
- Scaling-law extrapolation: distilled EOT-1B beats distilled Token-1B by up to 4% asymptotically, with far better data efficiency
The first large-scale empirical rebuttal of tokenizer necessity.
More from Research
- NULLs wins COLM Privacy & Security Workshop Best Paper for natively unlearnable LLMs — AdtRaghunathan · 2026-10-10
- Phantom Transfer: data poisoning survives 11 data-level defenses, NeurIPS 2026 paper shows — OwainEvans_UK · 2026-10-10
- Gym-Anything wins best paper at COLM 2026 Lifelong Agents Workshop — dan_fried · 2026-10-10
- DeepMind's Pushmeet Kohli: Why AlphaFold Didn't Actually Solve Protein Folding — Latent Space · 2026-10-10
- Ai2's Olmo Hybrid hits Olmo 3 7B MMLU accuracy with 49% fewer training tokens — allen_ai · 2026-10-10
- A 4B Model Trained for Under $500 Beats a 235B Sibling on Financial QA — AI Engineer · 2026-10-10