Pruning Q8 Models Could Match Q4 Size with Better Accuracy

An experiment on Kimi K2 suggests pruning experts from a Q8-quantized model could shrink it to the size of the current Q4 version (97GB) while retaining near-Q8 accuracy.

2026-08-27 ~ 2026-08-27 · 2 related posts