llama.cpp fails to load Qwen3.8 MTP draft model: 'output_hc_norm.weight' tensor not found
Ambitious_Fold_2874 · reddit · 2026-09-17
A Reddit user reports running unsloth's Qwen3.8-Flash-Next GGUF in llama.cpp works fine, but enabling MTP speculative decoding fails to load the draft model with error tensor 'outputhcnorm.weight' not found, causing llama-server to exit. Manually specifying --spec-type draft-mtp and the draft model path doesn't help; it's unclear whether the GGUF's MTP layers are incompatible with llama.cpp or the config is wrong.
More from Infra
- Hardware veteran: low-precision gains nearly exhausted, true sparsity is AI's next 10x — blelbach · 2026-09-17
- Moore's Law in three eras: from free lunch (1970-2005) to software hell (2015-now) — blelbach · 2026-09-17
- Fed's first rate hike in 3 years raises the financing bar for debt-funded GPU clusters — rohanpaul_ai · 2026-09-17
- GGUF Fit Calculator Reads File Headers to Tell Which Quants Fit Your GPU — asankhs · 2026-09-17
- What to run on a 96GB M3 Ultra? Qwen3.5-122B-A10B hits ~1000 tok/s locally — infieldmitt · 2026-09-17
- Prediction: Kimi K3-level AI on a single RTX 5090 within 18 months — TheZachMueller · 2026-09-17