Gemma 4 31B Quantization: Q4_K Draft Model Boosts Decode Speed by 10%
eightone-81 · reddit · 2026-08-07
A developer tested inference acceleration for the Gemma 4 31B model on dual 3090 GPUs. By quantizing the f16 MTP draft model to Q4K instead of the default Q40, they achieved a 10% decode speedup, increasing performance from 65 to 72 tokens per second.
The author also experimented with Q2K quantization, but the results degraded significantly, offering useful insights for local LLM quantization strategies.
More from Infra
- Tesla and SpaceX Joint Venture Plans 10M+ Sq Ft Terafab — scaling01 · 2026-08-07
- DeepSeek Resumes $8B Funding Round for Datacenter Buildout — scaling01 · 2026-08-07
- Running MiniMax H3 Video on RTX 3060: Turbo LoRA Enables 6-Step Generation — irmemon225 · 2026-08-07
- Cloudflare Launches Kitesurf: A Lightweight Browser Built for AI Agents — craigsdennis · 2026-08-07
- SpaceX and Tesla to Initially Spend $16.8 Billion on Terafab Chip Plant — pstAsiatech · 2026-08-07
- Together AI Demos Updated Inference Platform for Running Open Models in Production — togethercompute · 2026-08-07