Speculative Decoding Benchmarks: 36% Faster on Code, Losses on Prose
TheMoonMidas · x · 2026-09-02
Benchmarks for MTP and speculative decoding reveal that tok/s depends heavily on the content generated. Testing 6 engines on the same model and machine (M3 Max 96GB):
- Code & Math: Drafters perform well due to predictability.
- Prose: All drafters lose 17-36% performance as stories are unpredictable.
Results:
- mlx-serve: Fastest overall (Code 39.0, Math 41.7, Prose 32.1 tok/s).
- mlx-dspark & MTPLX: Math specialists.
- mlx-vlm: Best prose resilience (-17% loss).
Conclusion: A single tok/s figure is usually a coding or prose-specific metric; context matters.
More from Infra
- GLM-5.3 Model Gets GGUF Quantization Release for Edge Deployment — unsloth · 2026-09-02
- Asus AI PC Price Jumps 50%, Speculating on Upcoming DGX Spark Hike — mountainyoo · 2026-09-02
- User Praises GPT Infra Stability: Months Without Downtime — natesiggard · 2026-09-02
- Microsoft Research papers on LLM data infrastructure win awards at VLDB 2026 — jm_alexia · 2026-09-02
- Should you pay idle costs for local RAG just to keep batch jobs on the serving process? — Cautious_Bit_8521 · 2026-09-02
- M1 Max Benchmarks: 72 tok/s Aggregate Throughput at 128k Context — EyalToledano · 2026-09-02