Deep interview: why NVIDIA engineered Nemotron 3 Ultra around speed and long context

yacinelearning · x · 2026-10-06

A 1.5-hour interview with Chris, Senior Product Research Engineer on NVIDIA's Nemotron team, covering the engineering ethos behind the open Nemotron 3 Ultra: every design decision optimizes for speed and efficient long context. Topics include Latent MoE tradeoffs, aggressively grouped query attention, shared-weight MTP, MOPD, long context in-model vs. in-harness, the open model ecosystem, reward hacking stories, and advice for undergrads. Full timestamped TOC included.

Original post →

More from Infra

Infra channel →