Gemma 4 QAT Shows Significant Improvement in KV Cache Quantization, KLD Benchmarks Reveal

Anbeeld · reddit · 2026-08-12

The author used BeeLlama.cpp to compare the KV Cache quantization performance of Gemma 4 31B between standard quantization and QAT (Quantization-Aware Training). Benchmarks show that the QAT model exhibits a much stronger affinity for KV cache quantization.

Data indicates significant improvements across the board with QAT. For instance, at Q40 precision, QAT reduces KLD (KL Divergence) by nearly 10x while increasing Same-top agreement by about 15 percentage points. This proves that QAT can drastically reduce memory footprint while more effectively preserving the model's long-context processing capabilities.

Original post →

More from Models

Models channel →