Experiment: Pruning Q8 Quantized Models for Q4 Size with Higher Accuracy

EyalToledano · x · 2026-08-27

An experiment shared by the author suggests that pruning a Q8 quantized model (instead of a Q4 one) results in a post-prune size comparable to the current pre-prune Q4 model. This implies achieving Q8-level accuracy while maintaining the 97GB footprint of the current Q4 quantization.

Related event: Pruning Q8 Models Could Match Q4 Size with Better Accuracy(2 posts)→

Original post →

More from Infra

Infra channel →