Engineering Win: mxfp8 x mxfp4 Matmul Outperforms Standard mxfp8

zephyr_z9 · x · 2026-08-27

Discussion highlights that combining mxfp8 and mxfp4 for matrix multiplication yields higher performance than standard mxfp8 matmul. The technique involves 2x activation fetches plus concurrent scale sidecars to match fp4 native element sizes, presenting a non-trivial engineering challenge. Additionally, the Jalapeno chip defeated the VR200 on A0 silicon using a vibe-ported, non-speculative decoding implementation of DeepSeek, with faster program execution expected to improve further with a B0 update.

Original post →

More from Infra

Infra channel →