Pruning Q8 Models Could Match Q4 Size with Better Accuracy
An experiment on Kimi K2 suggests pruning experts from a Q8-quantized model could shrink it to the size of the current Q4 version (97GB) while retaining near-Q8 accuracy.
2026-08-27 ~ 2026-08-27 · 2 related posts
- Experiment: Pruning Q8 Quantized Models for Q4 Size with Higher Accuracy — EyalToledano · 2026-08-27
- Pruning Q8 experts aims for Q4 size with Q8 accuracy — EyalToledano · 2026-08-27