Paper: Speculative Decoding Often Slower on Mac, Best Case 1.61x

juanviera23 · reddit · 2026-08-15

A paper benchmarks speculative decoding on consumer hardware (Mac) and finds that 3 out of 5 configurations are slower than vanilla decoding, with the best case being 1.61x speedup at K=6. The author advises users to measure before enabling it. A Reddit user confirms similar results on llama.cpp and MLX.

Original post →

More from Infra

Infra channel →