Engineering Win: mxfp8 x mxfp4 Matmul Outperforms Standard mxfp8
zephyr_z9 · x · 2026-08-27
Discussion highlights that combining mxfp8 and mxfp4 for matrix multiplication yields higher performance than standard mxfp8 matmul. The technique involves 2x activation fetches plus concurrent scale sidecars to match fp4 native element sizes, presenting a non-trivial engineering challenge. Additionally, the Jalapeno chip defeated the VR200 on A0 silicon using a vibe-ported, non-speculative decoding implementation of DeepSeek, with faster program execution expected to improve further with a B0 update.
More from Infra
- Run GLM-5.3-Flash locally: 3-bit on 128GB RAM via Unsloth — StefanoGogioso · 2026-08-27
- Microsoft details Maia 200: no cache hierarchy, software-defined dataflow at 10K TFLOP/s FP4 — beenwrekt · 2026-08-27
- X Grok Bot Offers High-End Linux VM for Agents — vista8 · 2026-08-27
- Microsoft Maia 200: Eliminates Cache Hierarchy for Software Defined Dataflow — thoefler · 2026-08-27
- Cloudflare launches Computer: a virtual filesystem for AI agents — craigsdennis · 2026-08-27
- Pushing for LoRA sharing to reduce download waste — Borkato · 2026-08-27