Speculative Decoding Boosts 30B Model to 84 tok/s on a Single 24GB GPU
iam31337 · reddit · 2026-08-12
A deep dive into testing Muse Glimmer 30B with DFlash speculative decoding on a single RTX PRO 4000 (24GB), achieving up to 84.64 tok/s for code generation.
Key Findings & Optimizations:
- Acceptance Rate Misleads: Draft 4 had higher acceptance (47.46%) but lower speed (34.5 tok/s) compared to Draft 15 (17.99% acceptance, 47.20 tok/s).
- Backend Tuning: Moving draft argmax out of the CPU path and disabling llama.cpp's default min-p=0.05 removed synchronization bottlenecks.
- Quantization Trade-off: The fastest quant (NVFP4 hybrid at 94 tok/s) suffered worse perplexity; Q5KM proved to be the optimal complete system.
- 262K Context: Successfully decoded at 21.56 tok/s with 22.9GB VRAM usage when filled to the advertised 262K context limit.
The author concludes that under speculative decoding, raw tok/s or acceptance rates are misleading. A useful benchmark must account for workload distribution, exact quant recipes, draft lengths, and actual KV positions.
More from Infra
- FlashRT: AI Agents Auto-Optimize Multimodal Deployment, Cutting Latency by 70x — BeidiChen · 2026-08-12
- ComfyUI v0.32 Introduces New Native Attention, Matching SageAttention Speeds — slpreme · 2026-08-12
- Google Announces Three New Subsea Cables Connecting the Americas — rseroter · 2026-08-12
- DeepSeek V4 Quantization: Fixing Conversion Pitfalls and 8x RTX 5090 Benchmarks — gladkos · 2026-08-12
- Musk: Starlink to carry >90% of internet traffic, nearly 11K satellites in orbit — DimaZeniuk · 2026-08-12
- Apple Silicon Virtualization Breakthrough: LLM Speeds Up 16x on macOS VMs — Scobleizer · 2026-08-12