Ai2's Olmo Hybrid hits Olmo 3 7B MMLU accuracy with 49% fewer training tokens

allen_ai · x · 2026-10-10

Nonprofit lab Ai2 recapped its open research at COLM 2026 in San Francisco, spanning open models and AI for science agents.

The paper "Olmo Hybrid: From Theory to Practice and Back" combines transformer attention with linear recurrent layers: attention retrieves specific details from earlier text, while recurrent layers keep a compact running state. Controlled experiments show the hybrid reaches the same MMLU accuracy as Olmo 3 7B using 49% fewer training tokens.

The findings are informing the next Olmo model, now pre-training with a hybrid attention+recurrence mixture-of-experts architecture. Ai2 stresses that openly sharing models, data, code, and methods is core to its research approach.

Original post →

More from Research

Research channel →