Uno pairs AR weights with diffusion LoRA adapters, beating EAGLE-3 and all diffusion LLMs
HongyiWang10 · x · 2026-09-05
IFM released K2-Horizon-7B-Uno, a diffusion-augmented LLM that keeps the AR architecture but adds per-layer LoRA diffusion adapters enabling lossless parallel sampling. It outpaces all speculative decoding methods (DFlash, EAGLE-3) and beats every diffusion LLM (Mercury 2, Diffusion Gemma, Llada), scoring 70.1 on SWE-bench Verified and 93.0 on AIME-24 at 2.71 average TPF. Weights (LoRA adapter, Apache 2.0) are on Hugging Face.
Related event: Uno Architecture Combines Diffusion LLM Quality with AR Speed(4 posts)→
More from Research
- Princeton team says OpenAI's new architecture closely resembles its T2MLR paper — prfsanjeevarora · 2026-09-05
- Shared: article on recurrent-depth models and latent reasoning — ziv_ravid · 2026-09-05
- One policy controlling 4 different robot hands for in-hand manipulation rejected by CoRL — YuXiang_IRVL · 2026-09-05
- Orbis paper reframes video generation as a steerable, continuous visual process — rohanpaul_ai · 2026-09-05
- Training models to explain their own behavior: new interpretability dataset generalizes to held-out evals — a_karvonen · 2026-09-05
- LLMs produce inconsistent probabilities: P(rain) + P(no rain) don't sum to one — soumitrashukla9 · 2026-09-05