Intel Drops Fully MX-Compatible MXFP4/8 Quantized DeepSeek-V4-Flash Model
HaihaoShen · x · 2026-08-06
The Intel team has released a DeepSeek-V4-0731-Flash quantized model fully compatible with Microscaling (MX) formats. Optimized for native hardware inference, the key differences include: Dense Linears utilizing MXFP8 (compared to standard FP8), and MoE Linears combining MXFP4 weights with MXFP4 activations (compared to standard MXFP4 weights with FP8 activations).
More from Infra
- Tip: Limiting GPU Power Drastically Reduces Heat and Noise for Local AI — ForsakenAd1228 · 2026-08-06
- Recommended Free 'Inference Engineering' Book & Upcoming Book Club Kickoff — Al_Grigor · 2026-08-06
- x402 Standard Gains Linux Foundation Backing, AWS & Google Support for AI Agents — 0xJeff · 2026-08-06
- Set Up a Decentralized Private AI Inference Cluster with Bittensor — markjeffrey · 2026-08-06
- Elon Musk: 99% of Compute Will Be for AI Inference Long-Term — jamesdouma · 2026-08-06
- Hot Chips 2026 Opens Stanford Dorm Booking at $125/Night with $25 Student Grants — firstadopter · 2026-08-06