Nvidia Releases Compression-Optimized Nemotron Puzzle-75B
jacek2023 · reddit · 2026-07-07
Nvidia has released Nemotron-Labs-3-Puzzle-75B-A9B on Hugging Face, a new LLM optimized for deployment. It was compressed from Nemotron-3-Super-120B-A12B using an "iterative Puzzle" post-training compression framework. The model utilizes a hybrid MoE architecture, interleaving Mamba, MoE, and Attention layers, and supports Multi-Token Prediction (MTP). Following compression, total parameters dropped from 120.7B to 75.3B, and active parameters from 12.8B to 9.3B.
Nvidia claims that on an 8×B200 node, throughput is roughly 2x that of the original model, and concurrent 1M token processing on a single H100 increased from 1 to 8, while maintaining accuracy across reasoning, coding, multilingual, long-context, and agent benchmarks. It supports seven languages: English, French, German, Italian, Japanese, Spanish, and Chinese. It is now available for commercial use, with technical details available in the arXiv report.
Related event: NVIDIA Open-Sources 75B Parameter MoE Model Puzzle(4 posts)→
More from Models
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- awesome-llm-leaderboards: an open-source directory of LLM leaderboards, pricing tables, comparison tools — Last_Establishment_1 · 2026-09-11
- Anthropic claims it works to keep eval environments unidentifiable to models — MaxKannen · 2026-09-11
- Nex N2.5 Pro released on Hugging Face with 407GB of weights — jinnyjuice · 2026-09-11
- RoMa v2 image matching model unveiled in the usual black poster — ducha_aiki · 2026-09-11
- OpenAI rated Astra 'Critical' for cyber capabilities — and admits it's harder to monitor — theguywhobuilds · 2026-09-11