Meta extends FlashAttention-4 with MXFP8 on Blackwell, hitting 2.85 PFLOP/s forward

PyTorch · x · 2026-09-17

Meta Engineering extended FlashAttention-4 with MXFP8 support for NVIDIA Blackwell—forward and backward kernels, fused quantization, and jagged cross-attention—plus an end-to-end jagged module with fused quantization and FP8 activation/compute, already used internally for GEM training. On latest-gen hardware, the LP FA4 kernel reaches 2.85 PFLOP/s forward and 2 PFLOP/s backward, with up to 1.30× end-to-end module speedup. Design and open-source implementation detailed in their blog.

Original post →

More from Infra

Infra channel →