Uno Speculative Decoding Hits 1.30x on 28k-Token Prompts, 2x DFlash, Now in vLLM

yuntiandeng · x · 2026-10-02

Uno, a diffusion-augmented speculative decoding method from a paper by Eric Xing and colleagues, is being adopted by inference providers. On Gemma 4 26B with a 28k-token prompt on one RTX 3090, Uno delivers a lossless 1.30x speedup vs 0.66x for DFlash—2x faster—and beats plain decoding across 2k–28k lengths. Its drafter was trained on a single GPU with just 10k open prompts; Uno runs in vLLM today.

Original post →

More from Infra

Infra channel →