Uno: diffusion-augmented LLM matches AR quality with faster inference

Researchers from UIUC, Cornell, Cerebras and others propose Uno, a diffusion-augmented LLM that keeps the autoregressive architecture while adding diffusion, fixing quality and batch-inference speed gaps. An 8B Uno beats 26B DiffusionGemma and Mercury 2, with inference faster than EAGLE-3.

2026-09-04 ~ 2026-09-04 · 2 related posts