Speculative decoding boosts Qwen3.6-27B on one 5090, but slows crowded servers

luke_pacman · reddit · 2026-07-21

A detailed benchmark of Unsloth’s Qwen3.6-27B NVFP4 on one and two RTX 5090s shows that speculative decoding can help a lot — but only when the server is not already busy.

Key results:

The takeaway is that speculative decoding can be excellent for single-user latency, but once batching and concurrency saturate the GPU, it may hurt throughput instead of helping.

Original post →

More from Infra

Infra channel →