NVIDIA's Model-Optimizer unifies quantization, distillation, pruning and speculative decoding

NVIDIA · github · 2026-09-24

NVIDIA's open-source Model-Optimizer unifies SOTA compression techniques — quantization, distillation, pruning, NAS and speculative decoding — to compress models for deployment on TensorRT-LLM, TensorRT and vLLM for faster inference. The Python library has about 3.9k stars.

Original post →

More from Infra

Infra channel →