Uno Discrete Diffusion lands in K2 Horizon: faster than EAGLE-3, beats all diffusion LLMs
HongyiWang10 · x · 2026-09-08
The Kimi K2 Horizon release embeds "Uno Discrete Diffusion," with paper, models and code opened up so others can apply it to their own models.
- Uno targets the two known weaknesses of diffusion LLMs vs AR models: lower quality and slower inference at large batch sizes
- Method: keep the AR architecture, give each layer two sets of weights (AR and diffusion); the diffusion weights enable lossless parallel sampling from the AR distribution
- Results: faster than all speculative decoding methods (DFlash, EAGLE-3) and outperforms all diffusion LLMs (Mercury 2, Diffusion Gemma, Llada)
More from Models
- Desert Ant Labs launches: European lab building on-device intelligence — pcuenq · 2026-09-08
- A Codex session kept burning ~1,000 credits after its seat plan was switched mid-run — Liu_eroteme · 2026-09-08
- GPT-6 Astra vs MediaPipe for hand pose: SOTA reasoning, 3 min per frame — chris_j_paxton · 2026-09-08
- DeepSeek's Post-Sept 10 Pricing: Premium Output, Off-Peak Cache Hits at ¥0.02 — teortaxesTex · 2026-09-08
- Dev finds DayBreak model avoids Astra's cybersecurity safety triggers while testing extensions — HankYeomans · 2026-09-08
- Blind arena update: flash-next beats 3.5 max while Gemini 3.1 pro lags behind — Tall_Abrocoma_3533 · 2026-09-08