Pruning Q8 experts aims for Q4 size with Q8 accuracy
EyalToledano · x · 2026-08-27
The author discusses pruning experiments on the Kimi K2 model. The goal is to prune Q8 quantized experts so the model size matches the current Q4 version (97GB), effectively achieving near Q8 accuracy at the size of a Q4 model.
Related event: Pruning Q8 Models Could Match Q4 Size with Better Accuracy(2 posts)→
More from Infra
- ABF Substrates & PCB Identified as Key Constraints for 2027 AI Hardware — zephyr_z9 · 2026-08-27
- First startup enriches uranium for nuclear-powered data centers — Polymarket · 2026-08-27
- G2 spent $1.27M on 970B tokens, shifting focus from adoption to efficiency — prasanna_says · 2026-08-27
- Zhipu's GLM 3.5 Flash Served 42T Tokens in 6 Days Free on Chinese Chips — bindureddy · 2026-08-27
- TokenVisor supports Nvidia, AMD, and Intel GPUs in a single cluster — AccBalanced · 2026-08-27
- Zai's domestic inference cluster hits 100k+ chips; GLM-5.3 runs on custom interconnect — zephyr_z9 · 2026-08-27