Meta extends FlashAttention-4 with MXFP8 on Blackwell, hitting 2.85 PFLOP/s forward
PyTorch · x · 2026-09-17
Meta Engineering extended FlashAttention-4 with MXFP8 support for NVIDIA Blackwell—forward and backward kernels, fused quantization, and jagged cross-attention—plus an end-to-end jagged module with fused quantization and FP8 activation/compute, already used internally for GEM training. On latest-gen hardware, the LP FA4 kernel reaches 2.85 PFLOP/s forward and 2 PFLOP/s backward, with up to 1.30× end-to-end module speedup. Design and open-source implementation detailed in their blog.
More from Infra
- New dashboard sets first public baseline for advanced RL training costs — teortaxesTex · 2026-09-17
- PyTorch Conference NA lineup spotlights torch.compile and custom kernel breakthroughs — PyTorch · 2026-09-17
- Structured-decision trick speeds up DiffusionGemma inference 3-10x with one forward per request — bodonoghue85 · 2026-09-17
- Interactive guide maps the full AI compute stack from grid power to workloads — dr_alphalyrae · 2026-09-17
- How to run Qwen3.8-Flash-Next with N-gram SSD streaming in llama.cpp? — Ambitious_Fold_2874 · 2026-09-17
- How Bell Labs Missed the Microchip: IEEE Spectrum Revisits a Landmark Tech-History Blunder — ArtificialOther · 2026-09-17