AMD's official Qwen3.8 MXFP4 quant fails to load in llama.cpp, is support still missing?
mailto_devnull · reddit · 2026-09-03
Inspired by someone getting very high inference speeds on dual R9700s, a user tried AMD's official MXFP4 quant of Qwen3.8-27B (amd/Qwen3.8-27B-Quark-AWQ-MXFP4) on a single card, but its safetensors wouldn't load in llama.cpp. A community GGUF re-quant (magiccodingman's MagicQuant version) failed to load as well. The upshot: llama.cpp appears to lack MXFP4 support for now.
More from Infra
- fchollet posts a #TPU meme that leaves Hesamation exclaiming 'BRO WHAT?' — algo_diver · 2026-09-03
- Nvidia's CoWoS share seen falling to 45.6% by 2027 as AMD doubles to 19.8% — tengyanAI · 2026-09-03
- Loudoun County's 20-year data center history previews America's AI infrastructure future — suchenzang · 2026-09-03
- Alexandr Wang mocks NIMBYs: 'Planes come and go, but the data center hum never stops' — suchenzang · 2026-09-03
- GLM-5.3-Flash (320B MoE) on two DGX Sparks: rewritten CUDA kernel boosts fat-expert GEMM by ~40% — EAccelerate_42 · 2026-09-03
- Beff Jezos: Biology's Compute-per-Watt Is Massively Underestimated, Bio-Silicon Complexification Begins — beffjezos · 2026-09-03