Diffusion-augmented LLM Uno: 8B beats 26B DiffusionGemma with 3x lossless speedup
iScienceLuvr · x · 2026-09-04
A new paper (UIUC, Cornell, Cerebras; incl. Eric Xing) introduces diffusion-augmented LLMs: two weight sets — standard NTP-trained AR weights for quality, plus lightweight diffusion weights to draft multiple tokens in parallel via a Ψ-Spec sampler, giving lossless speedups.
Key results:
- The 8B Uno beats DFlash and Eagle3 speculative decoding on throughput at every batch size, without a separate draft model, with up to 3x speedup over the base AR model.
- 8B Uno outperforms the 26B open d-LLM DiffusionGemma and proprietary Mercury 2 across agentic tool use, coding, and long-context reasoning.
Paper, code, and HuggingFace weights are public.
Related event: Uno: diffusion-augmented LLM matches AR quality with faster inference(2 posts)→
More from Infra
- Extropic founder jokes Astra on Cerebras will feel like riding a Bugatti at 400kph — beffjezos · 2026-09-04
- DeepSeek V4 Flash Vision impresses via API, but 305B needs 4x GB300 to run — No_Issue_8224 · 2026-09-04
- Tim Sweeney says gaming faces worst crash since the 1980s as AI drives RAM costs — tekbog · 2026-09-04
- Transformers.js podcast: browser background removal, transcription and local agent workflows — nicodotdev · 2026-09-04
- Running Qwen3.8-Flash-Next with 256K context at 16 tok/s on DDR4 and a Tesla T4 — BusTiny207 · 2026-09-04
- Nvidia's PAIR turns your home network into a mini data center for local AI — The Decoder · 2026-09-04