Diffusion-augmented LLM Uno: 8B beats 26B DiffusionGemma with 3x lossless speedup

iScienceLuvr · x · 2026-09-04

A new paper (UIUC, Cornell, Cerebras; incl. Eric Xing) introduces diffusion-augmented LLMs: two weight sets — standard NTP-trained AR weights for quality, plus lightweight diffusion weights to draft multiple tokens in parallel via a Ψ-Spec sampler, giving lossless speedups.

Key results:

Paper, code, and HuggingFace weights are public.

Related event: Uno: diffusion-augmented LLM matches AR quality with faster inference(2 posts)→

Original post →

More from Infra

Infra channel →