Yacine's 90-minute deep dive: latent MoE, aggressive GQA inside Nvidia's open model

yacinelearning · x · 2026-10-01

Yacine released a 1h30 discussion with @llmwizard digging into the inner workings of NVIDIA's frontier-class open model.

Topics covered include latent MoE, speed optimizations, aggressive GQA configurations, and what he calls "bonkers" linearization of attention. The recurring thesis: a faster model is a smarter model — inference speed itself is capability.

Related event: Ex-NVIDIA Researcher Breaks Down Frontier Open Model Internals in 90-Minute Talk(2 posts)→

Original post →

More from Infra

Infra channel →