Qwen3.8 Distilled 9B/4B/2B Released: MMLU Scores Double
alejandroll10 · x · 2026-08-17
Community releases distilled models based on Qwen3.8-2.4T-A95B in 9B/4B/2B sizes. Via full-parameter SFT on curated teacher CoT, capabilities improved significantly: 9B MMLU rose from 54.6 to 75.1, 4B from 35.4 to 55.3, and 2B from 28.3 to 54.8. Weights and GGUF formats are available.
More from Models
- GLM-5.3 review: Cleaner code and consistent long-horizon performance — khademinori · 2026-08-17
- Seedance 2.5 Tops Multi-Image-to-Video Benchmark with Elo 1400 — rohanpaul_ai · 2026-08-17
- DeepSeek Adopts Peak-Valley Pricing: Off-Peak API at 50% Off — 智东西 · 2026-08-17
- Qwen2.5 2B runs on phones with 1GB RAM; MMLU jumps to 54.8 — alexcovo_eth · 2026-08-17
- New Claude/GPT APIs reportedly force temperature to 1 — andersonbcdefg · 2026-08-17
- What did Grok learn from watching this video? — Scobleizer · 2026-08-17