Qwen3.8-Flash-Next MTP lands in ik_llama.cpp mainline: 45 to 90 tok/s on RTX 5090, works on 12GB cards
Alternative_Will5974 · reddit · 2026-09-04
- The author's MTP support merged into ikllama.cpp mainline (PR #2369), no fork or patch needed
- The model's built-in 2.6B MTP head drafts tokens from hidden states; code draft acceptance hits 93-99%, prose only 60-65%
- Measured decode speeds: 45 → 90 tok/s on RTX 5090 + 128GB (experts on CPU); 85 → 113 on RTX Pro 6000 for code, but prose regressed 83 → 59
- 12GB 4070 got 9.5 → 12.5; a 3090 replicated it; head-target pairing matters, some combos go net negative
- Caveats: single-slot only (-np 1), --jinja hurts acceptance via default thinking mode
- Full build and server configs included, plus three HF repos with MTP head weights
More from Infra
- Micron explores near-GPU NAND flash to run bigger LLMs — giveen · 2026-09-04
- HPE Delivers Strong Q3 on AI Server Demand, Raises FY2026 Outlook — mattwbaker · 2026-09-04
- All Chromium Browsers Hit by Hard-to-Reproduce Data Loss Bug, Devs Say — uwukko · 2026-09-04
- Hundreds protest Scotland's datacentre boom, demanding pause on 20+ proposed projects — nordicinst · 2026-09-04
- Built a Dual RTX 6000 Pro Rig for Local DeepSeek — Warns Against Influencer Build Hype — HankYeomans · 2026-09-04
- Inference engines are an underexamined attack surface, self-hosting ops warned — JeremyCMorgan · 2026-09-04