Ai2's Olmo Hybrid hits Olmo 3 7B MMLU accuracy with 49% fewer training tokens
allen_ai · x · 2026-10-10
Nonprofit lab Ai2 recapped its open research at COLM 2026 in San Francisco, spanning open models and AI for science agents.
The paper "Olmo Hybrid: From Theory to Practice and Back" combines transformer attention with linear recurrent layers: attention retrieves specific details from earlier text, while recurrent layers keep a compact running state. Controlled experiments show the hybrid reaches the same MMLU accuracy as Olmo 3 7B using 49% fewer training tokens.
The findings are informing the next Olmo model, now pre-training with a hybrid attention+recurrence mixture-of-experts architecture. Ai2 stresses that openly sharing models, data, code, and methods is core to its research approach.
More from Research
- Jeremy Avigad's slides on the future of mathematics in the age of AI — ChengleiSi · 2026-10-10
- Tetris RL experiment: pretraining caps what RL can reach — PPO can provably converge to a bad policy — shizhediao · 2026-10-10
- Models trained on different data converge to similar concept geometry, researcher argues — cephaloform · 2026-10-10
- NULLs wins COLM Privacy & Security Workshop Best Paper for natively unlearnable LLMs — AdtRaghunathan · 2026-10-10
- Phantom Transfer: data poisoning survives 11 data-level defenses, NeurIPS 2026 paper shows — OwainEvans_UK · 2026-10-10
- DeepMind's Pushmeet Kohli on why AlphaFold didn't solve protein folding — Latent Space · 2026-10-10