Nvidia Releases Compression-Optimized Nemotron Puzzle-75B
jacek2023 · reddit · 2026-07-07
Nvidia has released Nemotron-Labs-3-Puzzle-75B-A9B on Hugging Face, a new LLM optimized for deployment. It was compressed from Nemotron-3-Super-120B-A12B using an "iterative Puzzle" post-training compression framework. The model utilizes a hybrid MoE architecture, interleaving Mamba, MoE, and Attention layers, and supports Multi-Token Prediction (MTP). Following compression, total parameters dropped from 120.7B to 75.3B, and active parameters from 12.8B to 9.3B.
Nvidia claims that on an 8×B200 node, throughput is roughly 2x that of the original model, and concurrent 1M token processing on a single H100 increased from 1 to 8, while maintaining accuracy across reasoning, coding, multilingual, long-context, and agent benchmarks. It supports seven languages: English, French, German, Italian, Japanese, Spanish, and Chinese. It is now available for commercial use, with technical details available in the arXiv report.
Related event: NVIDIA Open-Sources 75B Parameter MoE Model Puzzle(4 posts)→
More from Models
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11
- User Hails Unconfirmed 'DeepSeek 4.1 Flash' as an Inflection Point in LLMs — himanshustwts · 2026-09-11
- Terminal Bench v4: GLM-5.3 Leads at 41.9%, Kimi-K3 Underwhelms at 12.6% — Ok_Warning2146 · 2026-09-11
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11