Nvidia Releases Compression-Optimized Nemotron Puzzle-75B

jacek2023 · reddit · 2026-07-07

Nvidia has released Nemotron-Labs-3-Puzzle-75B-A9B on Hugging Face, a new LLM optimized for deployment. It was compressed from Nemotron-3-Super-120B-A12B using an "iterative Puzzle" post-training compression framework. The model utilizes a hybrid MoE architecture, interleaving Mamba, MoE, and Attention layers, and supports Multi-Token Prediction (MTP). Following compression, total parameters dropped from 120.7B to 75.3B, and active parameters from 12.8B to 9.3B.

Nvidia claims that on an 8×B200 node, throughput is roughly 2x that of the original model, and concurrent 1M token processing on a single H100 increased from 1 to 8, while maintaining accuracy across reasoning, coding, multilingual, long-context, and agent benchmarks. It supports seven languages: English, French, German, Italian, Japanese, Spanish, and Chinese. It is now available for commercial use, with technical details available in the arXiv report.

Related event: NVIDIA Open-Sources 75B Parameter MoE Model Puzzle(4 posts)→

Original post →

More from Models

Models channel →