SemiAnalysis breaks down NVIDIA LPU's three disaggregated inference configurations
SemiAnalysis_ · x · 2026-09-01
SemiAnalysis outlines three disaggregated (prefill/decode split) inference configurations supported by NVIDIA LPU:
- Rubin Prefill + LPU Decode: fastest interactivity;
- Rubin Prefill + Rubin Decode Attention + LPU Decode FFN: covers the middle of the latency curve;
- Rubin Prefill + Rubin Decode Verification + LPU Drafter: speculative decoding, for the middle-left of the curve.
For low-interactivity workloads, raw Rubin still wins. They look forward to Rubin + LPU performance curves on open-source agentic benchmarks like AgentX.
Related event: NVIDIA LPU Supports Three Disaggregated Inference Modes(2 posts)→
More from Infra
- TensorSharp vs llama.cpp: Qwen 3.8 Flash Next Benchmarks — fuzhongkai · 2026-09-01
- 2 engineers + AI designed a working LLM chip in 2 weeks, no human in the loop — 新智元 · 2026-09-01
- Why did increasing context size increase speed in Llama.cpp? — satnl · 2026-09-01
- AI inference demand surges again, supply brutally outpaced by token growth — Baconbrix · 2026-09-01
- Warp founder predicts cloud-based collaborative factories for all companies within a year — charlieholtz · 2026-09-01
- JPM: 1GW of AI Infrastructure Costs $40-45B, Frontier Labs Make ~$30B per GW — zephyr_z9 · 2026-09-01