W4A16 Quantization Boosts 30B Model Throughput by 4.5x on RTX 3090
mitchins-au · reddit · 2026-08-14
A developer released a W4A16 quantized version of NVIDIA's Nemotron 3.5 Lightning 30B-A3B model optimized for VLLM, allowing it to fit entirely on a single 24GB RTX 3090.
Tests show that W4A16 served via VLLM achieves several times higher throughput (around 4.5x at batch size 16) compared to the IQ4XS GGUF format on llama.cpp. Benchmarks indicate virtually no degradation in core capabilities like instruction following. This setup is ideal for consumer hardware handling batch labeling or agentic tasks where peak coding ability isn't critical.
More from Infra
- TPN Labs Announces Mainnet Competition to Tackle Edge AI Model Size Limits — const_reborn · 2026-08-14
- New Brain-Inspired AI Chip Solves Problems With 10,000x Fewer Calculations — ChuckDBrooks · 2026-08-14
- Architect Labs Uses AI to Design Custom Chips, Eliminating Need for In-House Semiconductor Teams — hsu_byron · 2026-08-14
- Qwen 30B MoE on RTX 3050 6GB: 30+ tps with 90k context — Bakkario · 2026-08-14
- Minimax with ref2va quantization runs on low VRAM — Actual-Project358 · 2026-08-14
- Continuous Batching in LLMs: The Tech Behind vLLM's 23x Throughput Jump — blaizedsouza · 2026-08-14