Diffusion-augmented LLM Uno delivers lossless 2.2x speedup over autoregressive generation

HongyiWang10 · x · 2026-09-18

IFMAI introduces Uno, a diffusion-augmented LLM aimed at the sequential token-by-token inference bottleneck of autoregressive models.

Key ideas:

Results: Uno can be trained from scratch or bolted onto open-weight AR LLMs; K2-Horizon-7B beats SOTA diffusion methods on both quality and throughput with up to 2.2x speedup, and outperforms leading speculative-decoding throughput at every evaluated batch size. Paper on arXiv (2609.04010), model released.

Original post →

More from Research

Research channel →