NVIDIA Paper: MoE Interactive Throughput Nearly Doubled

omarsar0 · x · 2026-07-08

omarsar0 highlights an NVIDIA paper on model compression. The work compresses the MoE model Nemotron-3-Super into Puzzle-75B-A9B, roughly doubling interactive serving throughput while preserving quality.

The core innovation is joint structural search: heterogeneous MoE pruning, active parameter budgets, and Mamba pruning are optimized simultaneously rather than sequentially. This is integrated into an iterative pipeline featuring distillation, reinforcement learning, quantization, and multi-token prediction heads. This joint optimization is key to the significant throughput boost under single-GPU interactive latency.

Related event: NVIDIA Open-Sources 75B Parameter MoE Model Puzzle(4 posts)→

Original post →

More from Infra

Infra channel →