Speculative Decoding Boosts RTX 5090 to 233 tok/s, Outperforming Mac
rohanpaul_ai · x · 2026-08-10
The author highlights that speculative decoding yields significantly higher performance gains on desktop GPUs compared to Macs. Benchmarks show inference speeds on an RTX-5090 surging from 74.9 to 233 tokens per second, while an M4-Max only moves from 23.7 to 38.
Speculative decoding uses a small, fast model to guess a block of tokens, which the main model then verifies in a single pass. The post specifically mentions Meta's DFlash, a lightweight drafter model shipped with Muse Glimmer, which proposes entire blocks of tokens at once for parallel verification, drastically improving inference efficiency.
More from Infra
- Discovered Materials Raises $9M to Hunt for Novel Chip Cooling Materials — TechCrunch AI · 2026-08-10
- 1M Token Context on Single RTX 3090 Achieved via KVarN Quantization — Anbeeld · 2026-08-10
- Choosing MiniMax H3 Quantization for RTX 5090: int8 vs nvfp4 — Zerozone000 · 2026-08-10
- MiniMax H3 Video Generation Stalls for 1 Hour on RTX 5090 — Johnwick1536 · 2026-08-10
- Offline KD Boosts Throughput 41% on Single H200, Slashes LLM Distillation Memory — MultiverseComputingCAI · 2026-08-10
- 50% higher costs: Why Chinese AI giants struggle to ditch Nvidia — pstAsiatech · 2026-08-10