Uno pairs AR weights with diffusion LoRA adapters, beating EAGLE-3 and all diffusion LLMs

HongyiWang10 · x · 2026-09-05

IFM released K2-Horizon-7B-Uno, a diffusion-augmented LLM that keeps the AR architecture but adds per-layer LoRA diffusion adapters enabling lossless parallel sampling. It outpaces all speculative decoding methods (DFlash, EAGLE-3) and beats every diffusion LLM (Mercury 2, Diffusion Gemma, Llada), scoring 70.1 on SWE-bench Verified and 93.0 on AIME-24 at 2.71 average TPF. Weights (LoRA adapter, Apache 2.0) are on Hugging Face.

Related event: Uno Architecture Combines Diffusion LLM Quality with AR Speed(4 posts)→

Original post →

More from Research

Research channel →