Building B200 Attention Kernel from Scratch to Near-SOTA in 60 Diagrams

dejavucoder · x · 2026-09-01

This blog post details building a dense B200 Attention Kernel from scratch using CUDA and PTX, reaching 94.4% of FlashAttention-4 performance, supported by 60 diagrams.

Key Content:

The post aims to provide a foundation for GPU kernel research on the latest hardware, noting that concepts apply beyond B200. Suitable for CUDA beginners and developers seeking deep optimization insights.

Original post →

More from Infra

Infra channel →