Auto-derived FlashAttention with SMEM and tensor core assignment shown off
vtabbott_ · x · 2026-09-12
vtabbott shares and hypes a demo of automatically derived FlashAttention, including SMEM management and tensor core assignment — meaning a compiler/search system can generate GPU attention kernels that previously required expert hand-tuning.
More from Infra
- Together serves 23%-30% of all OpenRouter traffic for GLM 5.3 models — zhyncs42 · 2026-09-12
- The Global Race for Cheap Power: Where AI Data Centers Should Actually Go — pravchaw · 2026-09-12
- VCs float 'hardware revenue derivative': fund compute costs via revenue share, not equity — ns123abc · 2026-09-12
- Relace hits 1T tokens/day on OpenRouter, serving 37% of DeepSeek v4 Flash traffic — stuffyokodraws · 2026-09-12
- Yutori's Navigator n2 runs browser agents at $1.46 per task on OSWorld 2.0 vs $13-$40+ for frontier models — DhruvBatra_ · 2026-09-12
- AI progress timing debate: same-node hardware gains deliver a one-time compute windfall — cis_female · 2026-09-12