Uno Speculative Decoding Hits 1.30x on 28k-Token Prompts, 2x DFlash, Now in vLLM
yuntiandeng · x · 2026-10-02
Uno, a diffusion-augmented speculative decoding method from a paper by Eric Xing and colleagues, is being adopted by inference providers. On Gemma 4 26B with a 28k-token prompt on one RTX 3090, Uno delivers a lossless 1.30x speedup vs 0.66x for DFlash—2x faster—and beats plain decoding across 2k–28k lengths. Its drafter was trained on a single GPU with just 10k open prompts; Uno runs in vLLM today.
More from Infra
- 180B Qwen model runs on one DGX Spark: 2.39-bit quant keeps 95.5% of BF16 scores — TheZachMueller · 2026-10-02
- Google: Starship must launch 1,600 times before space data centers work — TechCrunch AI · 2026-10-02
- Open-Source Local AI Avatar App Uses LM Studio and Chatterbox Voice Cloning — TheRedHairedHero · 2026-10-02
- Big Tech's AI Capex Now Accounts for Half of Wall Street's Profit Growth — speckx · 2026-10-02
- How to cost AI-powered filters: roofline model puts 5k-review LLM filter floor at 6.6s on H100 — sh_reya · 2026-10-02
- Hugging Face cofounder launches a million sandboxes live on stage at Modal Runtime — graceisford · 2026-10-02