NVIDIA LPU supports 3 disaggregated inference modes with Rubin architecture
teortaxesTex · x · 2026-09-01
NVIDIA LPU supports three types of disaggregated inferencing:
- Rubin Prefill + LPU Decode: For the fastest interactivity.
- Rubin Prefill + Rubin Decode Attention + LPU Decode FFN: For the middle of the performance curve.
- Rubin Prefill + Rubin Decode Verification + LPU Drafter: For the middle-left of the curve.
For low interactivity workloads, raw Rubin remains the winner. Performance on open-source agentic benchmarks like AgentX is anticipated.
Related event: NVIDIA LPU Supports Three Disaggregated Inference Modes(2 posts)→
More from Infra
- MTP released for Qwen3.8-Flash-Next GGUF, promising big local TPS gains — vini542reddit · 2026-09-01
- Google Cloud Monitoring MCP Connector Released — modelcontextprotocol · 2026-09-01
- Distributed.systems发布可审计的Agent基础设施 — arthurcolle · 2026-09-01
- Does enabling ChatGPT Memory or history reference increase token usage? — ssunki · 2026-09-01
- Engineer fixes ROCm inference crash on MI350X, uncovers 9 bugs in deep dive — AnushElangovan · 2026-09-01
- 40nm Neural-Dynamics Chip Uses Conductance Drift for 2.12ms Iteration Latency — maier_ak · 2026-09-01