Rubin Cuts Inference Cost Up to 10x, 2028 Kyber NVL1152 Will Link 1,152 GPUs in One NVLink Domain
imadade · reddit · 2026-09-10
A Reddit post maps NVIDIA's compute roadmap through 2028 and argues a new scaling axis is opening:
- 2025 Blackwell: the workhorse for current frontier training; latest internal models already mix Blackwell with early Rubin.
- 2026 Rubin: ramping into production as of August. NVIDIA claims 4x fewer GPUs for certain MoE training vs Blackwell and up to 10x lower inference token cost, with big HBM and interconnect gains that favor agent swarms and long reasoning traces. The author cites OpenAI's Navier-Stokes solution as an early sign.
- 2H 2027 Rubin Ultra: NVL576 topology links 576 GPUs across eight racks in one all-to-all NVLink domain (functional prototype already built on Blackwell), making hundreds of accelerators behave like one giant GPU and unlocking far more inference, search, verification and synthetic data.
- 2028 Feynman: confirmed at GTC 2026, built around die-stacked GPUs, next-gen custom HBM, Rosa CPU, LP40 inference hardware and BlueField-5. The centerpiece is NVLink 8 + Kyber NVL1152 — 1,152 GPUs across eight racks in one scale-up domain — targeting shared memory, inter-GPU bandwidth and inference throughput bottlenecks.
More from AGI Musings
- Runway CEO: Everyone's first instinct with a new frontier model is simulating reality — c_valenzuelab · 2026-09-10
- A first-year PhD student says ML systems research has been made obsolete by LLMs — kohjingyu · 2026-09-10
- AI slowdown talk is a cover for compute costs, argues viral critic — CtrlAltDwayne · 2026-09-10
- Jen Zhu Scott questions how the 'AI will kill us all' narrative took off without explanation — kevinnbass · 2026-09-10
- Tiny teams will look much bigger as AI gives them department-level output — alexmacgregor__ · 2026-09-10
- Models are overqualified too: frontier LLMs add endless complexity and cost to simple tasks — rachel_l_woods · 2026-09-10