DeepSeek Quantization Layer Pruning Experiment
mishig25 · x · 2026-07-12
Using DwarfStar with DeepSeek v4 Flash, the author tests the impact of applying better quantization methods to the "last 15% of layers."
Three metrics are used to measure quality:
- Avg NLL: Measures the model's "surprise" at the reference answer; lower is better.
- First-token matches: The proportion of times the model's most likely first token matches the reference answer; higher is better.
- Avg greedy LCP: The number of consecutively matched tokens before the first divergence; higher is better.
Experimentally, layers 37-42 (approx. 14%) represent antifreeze's default 2+4 hybrid model. Expanding this to layers 36-42 (approx. 16%) yields some additional gains, though the improvement is considered "relatively marginal." The author plans to continue researching.
More from Infra
- SkyPilot exits stealth with $20M seed round and an AI compute platform for fragmented clouds — jfiance · 2026-07-22
- Nothing phone mockup turns a film joke into a modular design meme — ZeYanjie · 2026-07-22
- Actual Computer says its inference stack is tuned for Nvidia’s consumer Blackwell lineup — markjeffrey · 2026-07-22
- Ben Bajarin says CPU demand is still being badly underestimated — BenBajarin · 2026-07-22
- An energy model says the U.S. could run short of natural gas starting in 2028 — churchkey · 2026-07-22
- Devin adds e2b sandboxes for remote agent execution — badphilosopher · 2026-07-22