Minima quantizes all 496 layers of Qwen3.8-27B to NVFP4 W4A4, matching BF16 at 2.9x smaller

pbaylies · x · 2026-09-05

A paper highlighted by HuggingPapers argues the recurrent half of a hybrid LLM is actually easy to quantize. Minima quantized all 496 linear layers of Qwen3.8-27B — including the Gated DeltaNet layers — to NVFP4 W4A4, matching BF16 performance while cutting model size by 2.9x.

Related event: Minima Quantizes 27B Hybrid LLM Fully to NVFP4 W4A4(2 posts)→

Original post →

More from Infra

Infra channel →