AMD's official Qwen3.8 MXFP4 quant fails to load in llama.cpp, is support still missing?

mailto_devnull · reddit · 2026-09-03

Inspired by someone getting very high inference speeds on dual R9700s, a user tried AMD's official MXFP4 quant of Qwen3.8-27B (amd/Qwen3.8-27B-Quark-AWQ-MXFP4) on a single card, but its safetensors wouldn't load in llama.cpp. A community GGUF re-quant (magiccodingman's MagicQuant version) failed to load as well. The upshot: llama.cpp appears to lack MXFP4 support for now.

Original post →

More from Infra

Infra channel →