GDN-2 variants substantially beat Mamba-2, GDP and KDA; 3B latent MoE release pending
teortaxesTex · x · 2026-09-26
Researcher ahatamiz1 (NVIDIA) reveals that larger variants of the new GDN-2 architecture have already been trained and substantially outperform competing approaches including Mamba-2, GDP, and KDA. A near-term release candidate is GDN-2-3B latent MoE, which follows the Nemotron-3 Nano architecture but replaces its Mamba-2 layers with GDN-2. A public release still requires several internal approvals, which the team is working to secure.
More from Research
- Contrastive World Models Learns World Models in Latent Space Without Pixel Prediction — burny_tech · 2026-09-27
- Rethinking on-policy distillation: researchers propose OLIVE, letting students learn from teacher continuations — May_F1_ · 2026-09-27
- CoRL 2026 workshop on continually self-improving robots opens call for papers, due Sep 28 — PeterStone_TX · 2026-09-27
- Martin Casado recommends the best talk on in-context learning, a first-principles view of LLMs — AccBalanced · 2026-09-27
- Functional Gradient Descent with Adaptive Representations accepted at NeurIPS — CatAstro_Piyush · 2026-09-27
- Tailored ASR for Japanese speaking assessment cuts mora error rate from 12.3% to 7.1% — tkasasagi · 2026-09-27