FULL STORY
Uno: Discrete Diffusion Brings Lossless Speedups to LLMs
An arXiv paper first proposed Uno, using discrete diffusion to overcome autoregressive inference bottlenecks. IFM AI then released the diffusion-augmented LLM, claiming lossless 2.2x speedups.
2026-09-08 ~ 2026-09-18 · 2 episodes · 6 posts
Episode 1 · Uno: Discrete Diffusion Delivers Lossless LLM Speedups (2026-09-08, 4 posts)
A new paper introduces Uno, which uses a lightweight diffusion adapter to draft tokens in parallel while the original autoregressive LLM verifies, achieving lossless 2.5x throughput gains on Qwen3-8B; authors also addressed overlap concerns with Orthrus.
- Uno: Diffusion-Augmented LLM Uses LoRA Drafts and Ψ-Spec Sampler for Lossless Speedups — burkov · 2026-09-08
- Uno speeds up Qwen3-8B 2.5x by using diffusion for parallel token drafting — rohanpaul_ai · 2026-09-09
- Uno paper: discrete diffusion drafting gives lossless LLM speedups without a draft model — rohanpaul_ai · 2026-09-09
- New discrete-diffusion LLM speedup paper Uno called out for similarity to Orthrus — _akhaliq · 2026-09-09
Episode 2 · IFM AI's Uno: Diffusion-Augmented LLM Delivers Lossless 2.2x Speedup (2026-09-18, 2 posts)
IFM AI released Uno, a diffusion-augmented LLM that tackles the token-by-token inference bottleneck by decoupling model parameters into two groups trained with standard NTP objectives, achieving a claimed lossless 2.2x speedup on K2 without degrading output quality.
- Diffusion-augmented LLM Uno delivers lossless 2.2x speedup over autoregressive generation — HongyiWang10 · 2026-09-18
- IFM AI's Uno couples diffusion with LLMs for lossless 2.2x faster generation — srchvrs · 2026-09-18