Minima quantizes all 496 layers of Qwen3.8-27B to NVFP4 W4A4, matching BF16 at 2.9x smaller
pbaylies · x · 2026-09-05
A paper highlighted by HuggingPapers argues the recurrent half of a hybrid LLM is actually easy to quantize. Minima quantized all 496 linear layers of Qwen3.8-27B — including the Gated DeltaNet layers — to NVFP4 W4A4, matching BF16 performance while cutting model size by 2.9x.
Related event: Minima Quantizes 27B Hybrid LLM Fully to NVFP4 W4A4(2 posts)→
More from Infra
- Dev Tests NVIDIA's Deepseek V4 NVFP4: 1M Context, 96% Memory at Batch 2048 — HankYeomans · 2026-09-05
- Open-Source On-Device Face Swap Hits Android: Real-Time on Hexagon NPU, 66 MB APK — Few_Caregiver8134 · 2026-09-05
- Who Needs a GB300? Researcher Makes a Movie on a $500 16GB GPU — francoisfleuret · 2026-09-05
- 19 latency patterns to cut non-model latency in AI applications — bibryam · 2026-09-05
- Astra reportedly trained on 100,000 GPUs, a staggering compute scale — sudoraohacker · 2026-09-05
- GPT-6 reportedly trained on just 100K B200/300 GPUs at a single Texas facility — round · 2026-09-05