IFM AI's Uno couples diffusion with LLMs for lossless 2.2x faster generation

srchvrs · x · 2026-09-18

IFM AI introduced Uno, a diffusion-augmented LLM claiming autoregressive quality at diffusion speed, targeting the token-by-token inference bottleneck. It's pitched as a lossless speedup that doesn't degrade response quality.

On K2-Horizon-7B, Uno reportedly beats state-of-the-art diffusion methods on both quality and throughput, with up to 2.2x speedup at no quality cost. Paper and model weights are public.

The reposter adds that Uno's plug-and-play design could extend to post-training and other complex scenarios, and argues diffusion scaling along both depth (denoising steps) and length (parallel test-time scaling) may unlock stronger reasoning.

Related event: IFM AI's Uno: Diffusion-Augmented LLM Delivers Lossless 2.2x Speedup(2 posts)→

Original post →

More from Models

Models channel →