SageAttention up to 1.96x faster causal attention on RX 9070 XT via hand-written HIP fp8 kernel

Familiar_Worry332 · reddit · 2026-10-02

A ComfyUI user released a drop-in SageAttention build for RX 9070/9070 XT on Windows, running a hand-written HIP fp8 attention kernel. Versus the existing gfx12 port (PR #368), attention calls are 1.07–1.21x faster non-causal and 1.72–1.96x faster causal; in ComfyUI with Krea2 that's 5–14% faster sampling per step than PyTorch SDPA. Keys: 64 keys per loop, transposed Q·Kᵀ keeping attention weights in registers, per-token scales instead of per-block — which avoids the 200x error blowups #368 sees on outlier-heavy layers like Krea2's first block. Caveats: gfx1201 only, Windows-tested, needs Python 3.12 + PyTorch 2.13.0+rocm10.0.0.

Related event: Hand-written HIP kernel brings SageAttention to RX 9070 XT with nearly 2x speedup(2 posts)→

Original post →

More from Infra

Infra channel →