NVIDIA's 75B hybrid MoE Nemotron-3-Puzzle is now runnable locally in llama.cpp
jacek2023 · reddit · 2026-09-03
- Model: NVIDIA Nemotron-3-Puzzle-75B-A9B can now run locally via llama.cpp.
- Architecture: Hybrid MoE with interleaved Mamba, MoE, and attention layers; supports Multi-Token Prediction (MTP) for faster generation.
- Specs: Compressed from parent Nemotron-3-Super-120B-A12B (120.7B total / 12.8B active) down to 75.3B total / 9.3B active parameters.
- Why it matters: An interesting size for local deployment — a drop-in option for those who couldn't run Nemotron-3-Super, and worth comparing against the newer Nemotron-3.5-Lightning-30B-A3B.
More from Infra
- Oxide's RFD 26 details why it picked bhyve and illumos over KVM and Xen — Sethwinterroth · 2026-09-03
- On-device LLM inference on Android with QNN, ExecuTorch, LiteRT and Gemma showcased at EdgeAI event — carrycooldude · 2026-09-03
- 20VC: Nvidia's $96.2B quarter, near-$12.9B Hugging Face deal, Cognition at $46BN — 20VC · 2026-09-03
- Polars 2.0 Pre-Release Announced: Big Upgrade for the Rust DataFrame Engine — komape · 2026-09-03
- Open-source ArtSmoker pipeline turns text prompts into textured, Blender-ready 3D models on your own AWS — niravdd · 2026-09-03
- One int8 op makes exact rescoring nearly free in PLAID-style late-interaction search — antoine_chaffin · 2026-09-03