Engineer tried 70B prefill + 7B decode back in 2023: combining spec dec with disagg was useless
suchenzang · x · 2026-09-03
A technical exchange between inference systems engineers: one asks why not combine speculative decoding AND disaggregated serving, to which @bingxu responds with first-hand experience — he tried exactly this kind of setup in 2023, using a 70B model for prefill and a 7B model for decode, and found it "entirely useless."
A useful data point for anyone exploring hybrid LLM serving architectures.
More from Infra
- KV cache, not parameter count, may be the real bottleneck for long-context local models — jonejy · 2026-09-03
- Developers warn a new Cloudflare feature could hurt your SEO — turn it off — gaganghotra_ · 2026-09-03
- If Amazon Trainium is any good, why weren't they at Hot Chips? — firstadopter · 2026-09-03
- Dragonfly rethinks Redis with sharded multi-threading to scale across modern multi-core servers — techNmak · 2026-09-03
- Zeiss exec: China is about 15 years behind in cutting-edge chipmaking tools — broodsugar · 2026-09-03
- Blogger argues Zeiss's mirror-coating know-how is only ~30 years deep, not magic — teortaxesTex · 2026-09-03