Qwen Pruning Tests: 256 Experts Optimal, Q4 Beats Q2
EyalToledano · x · 2026-08-27
Built a pruning curve for Qwen3.8-Flash-next to analyze size vs. accuracy tradeoffs:
- Sweet Spot: 256 experts is the optimal balance, halving size before performance buckles.
- Strategy Matters: Randomly cutting experts caused the model to collapse; selecting survivors is crucial.
- Quantization: Q2 quantization (full experts, 63GB) produced garbage output. Q4 quantization with half the experts performs significantly better.
More from Models
- Dry 5.6 Sol is sycophantic in chat but not when goal-oriented — Sauers_ · 2026-08-27
- Retrospective: Nous' First Model gpt4-x-vicuna-13b Trained on 180k GPT-4 Outputs — Teknium · 2026-08-27
- OpenAI Agents' Autonomy Analyzed; GLM-5.3 & Qwen4 Revealed Same Week — altryne · 2026-08-27
- Opus 5 uses 5x tokens vs GPT 5.6 for similar task accuracy — abeirami · 2026-08-27
- AdsBench Launches: Kimi K3 Tops AI Marketing Benchmark at $1.42/Task — qinzytech · 2026-08-27
- Zhipu's GLM 3.5 Flash Served 42T Tokens in 6 Days Free on Chinese Chips — bindureddy · 2026-08-27