Uno paper: discrete diffusion drafting gives lossless LLM speedups without a draft model
rohanpaul_ai · x · 2026-09-09
The arXiv paper "Unlocking Lossless Speedups in LLMs via Discrete Diffusion" introduces diffusion-augmented LLMs (Uno): model parameters are split into standard AR weights (trained with NTP) and lightweight diffusion weights (learned via a Diffusion Distillation phase with negligible overhead), enabling parallel multi-token generation.
- Ψ-Spec, a family of samplers, enables lossless acceleration and inference-time scaling at fixed context length
- Unlike speculative decoding, no separate draft model; unlike diffusion LLMs, quality of the underlying AR model is preserved
- Can be trained from scratch or added to existing open-weight AR LLMs; outperforms leading speculative-decoding throughput at every evaluated batch size
- Author team includes Eric Xing, Zhengzhong Liu, Joel Hestness, and others
Related event: Uno Uses Discrete Diffusion for Lossless LLM Speedups(3 posts)→
More from Infra
- MagicAILabs claims new recipe matches DeepSeek V4 Pro pretraining with 50x less compute, ~$0.5M — AccBalanced · 2026-09-09
- LM Studio ships Bionic 1.1.2 with faster sessions, better screen reader support, Linux builds — mattturck · 2026-09-09
- Navier-Stokes proof burned 300B output tokens, $20-30M at consumer API prices — soumitrashukla9 · 2026-09-09
- Nvidia and AMD fight over credit guarantees to bankroll customers' data centers — rohanpaul_ai · 2026-09-09
- DeepSeek 4 Flash runs all day on Spark: zero crashes, 98% cache hit, 40 tok/s — jasonkneen · 2026-09-09
- New report: custom ASIC market shifts from design wins to program responsibility — BenBajarin · 2026-09-09