Why Gated DeltaNet survives 4-bit: full NVFP4 W4A4 quantization of a hybrid 27B LLM
minima-ai · hf · 2026-09-04
minima-ai fully quantized a hybrid 27B LLM — including its recurrent Gated DeltaNet layers — to NVFP4 W4A4, preserving accuracy on long-context and reasoning benchmarks. The key is localizing outliers and exploiting the robust delta-rule dynamics, suggesting recurrent halves of hybrid LLMs are naturally low-bit friendly for edge deployment.
More from Infra
- Nvidia crushes quarter and buys Hugging Face; Cognition raises at $46BN, Linear $2.5BN — IanAndrewsDC · 2026-09-04
- KV Cache Scoring Is Useless: Random Eviction Matches Top Evictors at 32-43% Higher Throughput — burny_tech · 2026-09-04
- Commenter: Anyone against data centers in 2026 is fundamentally unserious — RachelVT42 · 2026-09-04
- 16k runs reveal which tools Claude Code, Codex and Cursor actually pick — Saboo_Shubham_ · 2026-09-04
- Reuters: Nvidia to buy Hugging Face for nearly $13B in big bet on open AI models — Chief_Taquero · 2026-09-04
- A 4-year-old RTX 4080 now costs $1,500, and the AI community is laughing at the pricing — yacineMTB · 2026-09-04