Uno speeds up Qwen3-8B 2.5x by using diffusion for parallel token drafting
rohanpaul_ai · x · 2026-09-09
Uno introduces a way to speed up existing LLMs without changing their output distribution: keep the original autoregressive model in charge of quality, and use lightweight diffusion adapters to draft multiple tokens in parallel, which the base model then verifies. This removes the need for a separate draft model and preserves the base model's sampling behavior.
- On Qwen3-8B, Uno delivered 2.5x higher per-request throughput and 1.6x higher system throughput at the largest tested batch size
- End-to-end RL training sped up by up to 40% in reported runs
Related event: Uno Uses Discrete Diffusion for Lossless LLM Speedups(3 posts)→
More from Infra
- SageMaker Feature Store adds UpdateRecord for atomic feature-level writes, no more read-modify-write — AWS ML Blog · 2026-09-09
- LM Studio ships Bionic 1.1.2 with faster sessions, better screen reader support, Linux builds — mattturck · 2026-09-09
- Navier-Stokes proof burned 300B output tokens, $20-30M at consumer API prices — soumitrashukla9 · 2026-09-09
- Nvidia and AMD fight over credit guarantees to bankroll customers' data centers — rohanpaul_ai · 2026-09-09
- DeepSeek 4 Flash runs all day on Spark: zero crashes, 98% cache hit, 40 tok/s — jasonkneen · 2026-09-09
- New report: custom ASIC market shifts from design wins to program responsibility — BenBajarin · 2026-09-09