Nemotron 3 Ultra's 550B params mogged by models 20x smaller, SemiAnalysis blasts NVIDIA's frontier training

teortaxesTex · x · 2026-09-03

SemiAnalysis argues that while NVIDIA Research produces excellent architecture work (LatentMoE used in Kimi K3, GatedDeltaNets in Qwen), NVIDIA's bureaucratic culture undermines end-to-end frontier training: Nemotron 3 Ultra, with 550B total params (55B active), is outperformed by all Chinese models, including Qwen3.8 27B with 20x fewer parameters. teortaxesTex adds that architecture R&D output tends to be anticorrelated with model quality and questions the dogged criticism of Nemotron 3.

Original post →

More from Companies & People

Companies & People channel →