Gemma 4 NVFP4 Quantization Released
kalyan_kpl · x · 2026-07-15
Unsloth released the NVFP4 quantized version of Gemma 4, focusing on faster GPU inference and lower VRAM usage.
- The 12B version can run on 11GB VRAM
- The 26B-A4B achieves 13K tok/s on the B200
- Official claims show a 1.5× inference speedup compared to existing solutions
- The focus is on more accurate 4-bit inference optimization tailored for the Blackwell architecture
The post also includes blog and model links, representing a concrete and usable inference optimization update.
Related event: Unsloth Releases NVFP4 Quantized Gemma 4 for Faster Inference(2 posts)→
More from Infra
- Devin adds e2b sandboxes for remote agent execution — badphilosopher · 2026-07-22
- Arbitrum fee simulation shows higher gas capacity but lower L2 revenue under ArbOS61 — tomwanhh · 2026-07-22
- NVIDIA pushes OpenUSD as the common layer for simulation and physical AI — MonaJalal_ · 2026-07-22
- SkyPilot exits stealth with $20M to unify fragmented GPU compute across five clouds — skypilot_org · 2026-07-22
- Production AI budgets include retries, routing, caching and observability—not just token prices — arx-go · 2026-07-22
- NVIDIA briefs analysts on Vera CPU and doubles down on monolithic agentic design — BenBajarin · 2026-07-22