W4A16 Quantization Boosts 30B Model Throughput by 4.5x on RTX 3090

mitchins-au · reddit · 2026-08-14

A developer released a W4A16 quantized version of NVIDIA's Nemotron 3.5 Lightning 30B-A3B model optimized for VLLM, allowing it to fit entirely on a single 24GB RTX 3090.

Tests show that W4A16 served via VLLM achieves several times higher throughput (around 4.5x at batch size 16) compared to the IQ4XS GGUF format on llama.cpp. Benchmarks indicate virtually no degradation in core capabilities like instruction following. This setup is ideal for consumer hardware handling batch labeling or agentic tasks where peak coding ability isn't critical.

Original post →

More from Infra

Infra channel →