Yacine's 90-minute deep dive: latent MoE, aggressive GQA inside Nvidia's open model
yacinelearning · x · 2026-10-01
Yacine released a 1h30 discussion with @llmwizard digging into the inner workings of NVIDIA's frontier-class open model.
Topics covered include latent MoE, speed optimizations, aggressive GQA configurations, and what he calls "bonkers" linearization of attention. The recurring thesis: a faster model is a smarter model — inference speed itself is capability.
More from Infra
- Ai2 Releases Olmo-core 3: Open MoE Training Framework Scales to 128 Experts with <5% Throughput Loss — allen_ai · 2026-10-01
- Ai2's tech report shows how 512 GPUs train one model—and flags 'token gerrymandering' — allen_ai · 2026-10-01
- Ai2 open-sources Olmo-core 3, training infrastructure scaling MoE models toward trillion parameters — allen_ai · 2026-10-01
- Dwarfstar's Bespoke Quants Run Qwen Fast on a 96GB M3 Ultra — TheRealJesus2 · 2026-10-01
- Micron beat the call but the model was $1.9B off: an analyst's post-earnings postmortem — tengyanAI · 2026-10-01
- Micron's $54B quarter pulls the whole HBM packaging, probe, and inspection capacity chain — demian_ai · 2026-10-01