PyTorch's TLX-based JFA kernel beats FlashAttention-4 by 13% fwd, 50% bwd on B200
PyTorch · x · 2026-10-02
PyTorch published a blog detailing Jagged Flash Attention (JFA), the attention kernel behind Meta's Generative Ads Model (GEM), running on NVIDIA Blackwell B200.
- Built with TLX (Triton Low-level Extensions), which add explicit, hardware-aware control on top of Triton's tile-based programming model.
- 3.2K lines of concise Triton-level code — roughly 3× less than the 10K-line CuteDSL kernels of FlashAttention-4 (FA4).
- Outperforms FA4 on the jagged shapes that matter for GEM: 13% faster on forward pass and 50% faster on backward pass.
The post is a strong case for writing near-hand-tuned kernels in a high-level language, with direct relevance for inference/training kernel engineers.
More from Infra
- Lambda closes $1B-plus GPU debt financing at 6.78% fixed rate, investment-grade rated — TheZachMueller · 2026-10-02
- Dual DGX Spark cluster vs M5 Ultra Mac Studio benchmark video is out — AIFlow_ML · 2026-10-02
- SageAttention up to 1.96x faster causal attention on RX 9070 XT via hand-written HIP fp8 kernel — Familiar_Worry332 · 2026-10-02
- Local GLM setup reportedly hits 175 tok/s with 98% draft acceptance on dual RTX 6000 Pros — HankYeomans · 2026-10-02
- Cloudflare lets non-admin members self-serve create Account API tokens via CLI, API and Terraform — irvinebroque · 2026-10-02
- SpaceX launches Google AI chips into orbit in push toward space-based data centers — pstAsiatech · 2026-10-02