Gemma 4 Quantization: Recovering 96% Reasoning Performance via Precision Redistribution
devildip · reddit · 2026-08-15
A study demonstrates the dramatic impact of Tensor Level Quantization Allocation under extreme budget constraints.
Key Results:
- At a 3.3GB IQ2XXS limit, reasoning scores improved from 28.9 to 69.5 (+140.54%).
- The final model is only 24% the size of the BF16 source but retains 96.74% of its reasoning performance.
- 10 out of 11 evaluated categories outperformed the imatrix baseline, with only stability regressing.
Methodology:
- No post-training, LoRA, or pruning involved.
- Builds an imatrix from category-oriented corpora, measures damage, and redistributes precision at the tensor level under a fixed byte budget.
Capability Retention:
- Knowledge QA: 97.50%
- Context: 95.83%
- Instruction following: 81.25%
- Coding: 58.49%
Related event: Tensor-Level Quantization Keeps Gemma 4 Strong at Tiny Sizes(2 posts)→
More from Models
- User Feedback: Qwen3.8 Overthinks and Misses the Point on Tasks — WhatererBlah555 · 2026-08-15
- Qwen3.8 27B评测:一次性编程能力显著提升 — kms_dev · 2026-08-15
- Whisper: Local speech-to-text system for multiple languages — ZabihullahAtal · 2026-08-15
- Fable is top but token-heavy; GPT models more efficient with similar intelligence — gethackteam · 2026-08-15
- Qwen3.8-35B-A3B Spotted in GitHub Code, Optimized for Consumer GPUs — cephaloform · 2026-08-15
- MiniMax H3 Re2vA Release — smereces · 2026-08-15