Intel Drops Fully MX-Compatible MXFP4/8 Quantized DeepSeek-V4-Flash Model

HaihaoShen · x · 2026-08-06

The Intel team has released a DeepSeek-V4-0731-Flash quantized model fully compatible with Microscaling (MX) formats. Optimized for native hardware inference, the key differences include: Dense Linears utilizing MXFP8 (compared to standard FP8), and MoE Linears combining MXFP4 weights with MXFP4 activations (compared to standard MXFP4 weights with FP8 activations).

Original post →

More from Infra

Infra channel →