Auto-derived FlashAttention with SMEM and tensor core assignment shown off

vtabbott_ · x · 2026-09-12

vtabbott shares and hypes a demo of automatically derived FlashAttention, including SMEM management and tensor core assignment — meaning a compiler/search system can generate GPU attention kernels that previously required expert hand-tuning.

Original post →

More from Infra

Infra channel →