Pruning Q8 experts aims for Q4 size with Q8 accuracy

EyalToledano · x · 2026-08-27

The author discusses pruning experiments on the Kimi K2 model. The goal is to prune Q8 quantized experts so the model size matches the current Q4 version (97GB), effectively achieving near Q8 accuracy at the size of a Q4 model.

Related event: Pruning Q8 Models Could Match Q4 Size with Better Accuracy(2 posts)→

Original post →

More from Infra

Infra channel →