llama.cpp fails to load Qwen3.8 MTP draft model: 'output_hc_norm.weight' tensor not found

Ambitious_Fold_2874 · reddit · 2026-09-17

A Reddit user reports running unsloth's Qwen3.8-Flash-Next GGUF in llama.cpp works fine, but enabling MTP speculative decoding fails to load the draft model with error tensor 'outputhcnorm.weight' not found, causing llama-server to exit. Manually specifying --spec-type draft-mtp and the draft model path doesn't help; it's unclear whether the GGUF's MTP layers are incompatible with llama.cpp or the config is wrong.

Original post →

More from Infra

Infra channel →