Engineer tried 70B prefill + 7B decode back in 2023: combining spec dec with disagg was useless

suchenzang · x · 2026-09-03

A technical exchange between inference systems engineers: one asks why not combine speculative decoding AND disaggregated serving, to which @bingxu responds with first-hand experience — he tried exactly this kind of setup in 2023, using a 70B model for prefill and a 7B model for decode, and found it "entirely useless."

A useful data point for anyone exploring hybrid LLM serving architectures.

Related event: Heterogeneous prefill-decode split sparks debate after engineer reveals 2023 trial(3 posts)→

Original post →

More from Infra

Infra channel →