Will Nvidia Vera Rubin Actually Speed Up LLM Pre-Training? And Are 10T+ Models Next?

Witty_County5128 · reddit · 2026-09-29

A user asks whether Nvidia's Vera Rubin numbers, which look huge, mostly reflect low-precision formats and inference, and how much of that carries over to pre-training.

He also wonders whether 10T+ parameter models are coming, or whether data, power and cost are now the real limits, keeping the focus on MoE and better data rather than sheer scale. He invites input from people with hardware or training experience.

Original post →

More from Infra

Infra channel →