Uno speeds up Qwen3-8B 2.5x by using diffusion for parallel token drafting

rohanpaul_ai · x · 2026-09-09

Uno introduces a way to speed up existing LLMs without changing their output distribution: keep the original autoregressive model in charge of quality, and use lightweight diffusion adapters to draft multiple tokens in parallel, which the base model then verifies. This removes the need for a separate draft model and preserves the base model's sampling behavior.

Related event: Uno Uses Discrete Diffusion for Lossless LLM Speedups(3 posts)→

Original post →

More from Infra

Infra channel →