Deep interview: why NVIDIA engineered Nemotron 3 Ultra around speed and long context
yacinelearning · x · 2026-10-06
A 1.5-hour interview with Chris, Senior Product Research Engineer on NVIDIA's Nemotron team, covering the engineering ethos behind the open Nemotron 3 Ultra: every design decision optimizes for speed and efficient long context. Topics include Latent MoE tradeoffs, aggressively grouped query attention, shared-weight MTP, MOPD, long context in-model vs. in-harness, the open model ecosystem, reward hacking stories, and advice for undergrads. Full timestamped TOC included.
More from Infra
- Pi-hole-class DNS ad-blocker runs on a $2 ESP32-C3 with 537k domains in flash — M-Abozaid · 2026-10-06
- Cloudflare Lets Workers Connect to Artifacts Repos, Cutting GitHub Out of the Build Pipeline — threepointone · 2026-10-06
- Vultr books $1.2B AMD AI rack order as buyers reserve capacity years ahead — shashib · 2026-10-06
- Swapping AdamW States for FFT Cuts Fine-tuning VRAM by 50% Without Quantization — Spectra-Global · 2026-10-06
- NanoGPT speedrun sets record: 11.3% faster via architecture-only change, paper coming — yoavartzi · 2026-10-06
- NVIDIA's CANTO Predicts Aerodynamics Directly From CAD, Cuts Pressure Error 20% — JeanKossaifi · 2026-10-06