Gemma 4 31B Quantization: Q4_K Draft Model Boosts Decode Speed by 10%

eightone-81 · reddit · 2026-08-07

A developer tested inference acceleration for the Gemma 4 31B model on dual 3090 GPUs. By quantizing the f16 MTP draft model to Q4K instead of the default Q40, they achieved a 10% decode speedup, increasing performance from 65 to 72 tokens per second.

The author also experimented with Q2K quantization, but the results degraded significantly, offering useful insights for local LLM quantization strategies.

Original post →

More from Infra

Infra channel →