Nvidia Releases Compression-Optimized Nemotron Puzzle-75B
jacek2023 · reddit · 2026-07-07
Nvidia has released Nemotron-Labs-3-Puzzle-75B-A9B on Hugging Face, a new LLM optimized for deployment. It was compressed from Nemotron-3-Super-120B-A12B using an "iterative Puzzle" post-training compression framework. The model utilizes a hybrid MoE architecture, interleaving Mamba, MoE, and Attention layers, and supports Multi-Token Prediction (MTP). Following compression, total parameters dropped from 120.7B to 75.3B, and active parameters from 12.8B to 9.3B.
Nvidia claims that on an 8×B200 node, throughput is roughly 2x that of the original model, and concurrent 1M token processing on a single H100 increased from 1 to 8, while maintaining accuracy across reasoning, coding, multilingual, long-context, and agent benchmarks. It supports seven languages: English, French, German, Italian, Japanese, Spanish, and Chinese. It is now available for commercial use, with technical details available in the arXiv report.
Related event: NVIDIA Open-Sources 75B Parameter MoE Model Puzzle(4 posts)→
More from Models
- Google says information agents are coming to AI Pro and Ultra this summer — gaganghotra_ · 2026-07-22
- Poolside’s Laguna S 2.1 gets a two-week free run on Nous Portal — NousResearch · 2026-07-22
- Qwen3.8 Max Preview looks substantially better in a side-by-side test with Kimi K3 — curiousily_ · 2026-07-22
- Moonshot’s Kimi K3 reaches #5 on MathArena as the top open model — xeophon · 2026-07-22
- Google launches Gemini 3.5 Flash Cyber for CodeMender, with limited access for governments — GoogleAI · 2026-07-22
- Gemini 3.5 Flash-Lite beats 3.1 Flash-Lite on long-context retrieval in MRCRv2 — Dillonu · 2026-07-22