SageAttention up to 1.96x faster causal attention on RX 9070 XT via hand-written HIP fp8 kernel
Familiar_Worry332 · reddit · 2026-10-02
A ComfyUI user released a drop-in SageAttention build for RX 9070/9070 XT on Windows, running a hand-written HIP fp8 attention kernel. Versus the existing gfx12 port (PR #368), attention calls are 1.07–1.21x faster non-causal and 1.72–1.96x faster causal; in ComfyUI with Krea2 that's 5–14% faster sampling per step than PyTorch SDPA. Keys: 64 keys per loop, transposed Q·Kᵀ keeping attention weights in registers, per-token scales instead of per-block — which avoids the 200x error blowups #368 sees on outlier-heavy layers like Krea2's first block. Caveats: gfx1201 only, Windows-tested, needs Python 3.12 + PyTorch 2.13.0+rocm10.0.0.
More from Infra
- Modal launches BYOC 2.0 to keep cloud coding agents' sensitive data inside your account — charles_irl · 2026-10-02
- GPT-6 Astra Ultrafast launches on NVIDIA Blackwell with up to 8x faster tokens — nvidia · 2026-10-02
- Google's Project Suncatcher prototype satellite is now in orbit — Gaiden206 · 2026-10-02
- OpenAI's GPT-6 Astra Ultrafast Delivers Up to 8x Faster Token Generation on NVIDIA Blackwell — NVIDIA Blog · 2026-10-02
- Ben Lorica: The AI data problem moved downstream, from finding data to making it usable — bigdata · 2026-10-02
- California man arrested for allegedly smuggling over $300M in AI servers to China — Polymarket · 2026-10-02