PyTorch's TLX-based JFA kernel beats FlashAttention-4 by 13% fwd, 50% bwd on B200

PyTorch · x · 2026-10-02

PyTorch published a blog detailing Jagged Flash Attention (JFA), the attention kernel behind Meta's Generative Ads Model (GEM), running on NVIDIA Blackwell B200.

The post is a strong case for writing near-hand-tuned kernels in a high-level language, with direct relevance for inference/training kernel engineers.

Original post →

More from Infra

Infra channel →