Will Nvidia Vera Rubin Actually Speed Up LLM Pre-Training? And Are 10T+ Models Next?
Witty_County5128 · reddit · 2026-09-29
A user asks whether Nvidia's Vera Rubin numbers, which look huge, mostly reflect low-precision formats and inference, and how much of that carries over to pre-training.
He also wonders whether 10T+ parameter models are coming, or whether data, power and cost are now the real limits, keeping the focus on MoE and better data rather than sheer scale. He invites input from people with hardware or training experience.
More from Infra
- Stanford CS153 recap: Google depreciates compute over six years, NVIDIA ships a new arch every year — le_james94 · 2026-09-29
- Huang's compute logic: Amdahl's Law caps speedups, over-provision the slow parts — le_james94 · 2026-09-29
- Enrichment bottlenecks nuclear, nuclear bottlenecks power, power bottlenecks AI — le_james94 · 2026-09-29
- Turso 0.8.0 lands: up to 7x faster than SQLite with 500x lower tail latency — glcst · 2026-09-29
- AMD ships Ryzen AI Max+ PRO with 192GB unified memory, 50% more than NVIDIA's upcoming Spark — ryanshrout · 2026-09-29
- Founders now spend their time hunting compute: Instinct demand doubles every week — ai · 2026-09-29