NVIDIA LPU Supports Three Disaggregated Inference Modes

SemiAnalysis details three disaggregated (prefill/decode split) inference configurations supported by NVIDIA's LPU, paired with the Rubin architecture, with Rubin Prefill + LPU Decode offering the fastest interactive performance.

2026-09-01 ~ 2026-09-01 · 2 related posts