NVIDIA Updates 30B Hybrid Model: Distillation Delivers 70% Intelligence Boost

PavloMolchanov · x · 2026-08-11

NVIDIA released a weight update for its smallest 30B-A3B model (MoE+Mamba architecture) from the NVIDIA-Nemotron family. By distilling knowledge from its larger models (Super and Ultra) during continuous pre-training and post-training, the model's capability has significantly improved.

The official AA index jumped from 14 to 24, achieving a 70% performance gain without any architectural changes. This means users can get substantially more intelligence simply by updating the weight file. The new score closely approaches their previous Super model, which is 4x larger. The architecture combines Mamba2 for efficient long-context handling and MoE to save forward compute while maintaining intelligence.

Original post →

More from Models

Models channel →