Speculative Decoding Benchmarks: 36% Faster on Code, Losses on Prose

TheMoonMidas · x · 2026-09-02

Benchmarks for MTP and speculative decoding reveal that tok/s depends heavily on the content generated. Testing 6 engines on the same model and machine (M3 Max 96GB):

Results:

Conclusion: A single tok/s figure is usually a coding or prose-specific metric; context matters.

Original post →

More from Infra

Infra channel →