Nemotron 3 Ultra's 550B params mogged by models 20x smaller, SemiAnalysis blasts NVIDIA's frontier training
teortaxesTex · x · 2026-09-03
SemiAnalysis argues that while NVIDIA Research produces excellent architecture work (LatentMoE used in Kimi K3, GatedDeltaNets in Qwen), NVIDIA's bureaucratic culture undermines end-to-end frontier training: Nemotron 3 Ultra, with 550B total params (55B active), is outperformed by all Chinese models, including Qwen3.8 27B with 20x fewer parameters. teortaxesTex adds that architecture R&D output tends to be anticorrelated with model quality and questions the dogged criticism of Nemotron 3.
More from Companies & People
- Alexandr Wang mocks NIMBYs: 'Planes come and go, but the data center hum never stops' — suchenzang · 2026-09-03
- OpenAI loses safety leadership: ethics, safety systems and mission alignment heads all exit as preparedness team is restructured — austinc3301 · 2026-09-03
- Meta Ends AI Usage Quotas in Performance Reviews While Rolling Out Agent Tool Hatch — nordicinst · 2026-09-03
- [un]prompted.au AI security conference sells out, adds virtual tickets for Sept 2026 Sydney event — moyix · 2026-09-03
- AI x cybersecurity conference [un]prompted.au reveals 24-session program for Sydney 2026 — dyn___ · 2026-09-03
- Matthew Chang's team engineered $500M of US factory automation in 12 months, with $1B more queued — MatthewChang · 2026-09-03