Speculative Decoding Boosts RTX 5090 to 233 tok/s, Outperforming Mac

rohanpaul_ai · x · 2026-08-10

The author highlights that speculative decoding yields significantly higher performance gains on desktop GPUs compared to Macs. Benchmarks show inference speeds on an RTX-5090 surging from 74.9 to 233 tokens per second, while an M4-Max only moves from 23.7 to 38.

Speculative decoding uses a small, fast model to guess a block of tokens, which the main model then verifies in a single pass. The post specifically mentions Meta's DFlash, a lightweight drafter model shipped with Muse Glimmer, which proposes entire blocks of tokens at once for parallel verification, drastically improving inference efficiency.

Original post →

More from Infra

Infra channel →