Hand-written Blackwell GEMM in CuTe DSL hits 1401 TFLOP/s, about 97% of cuBLAS on B200
retr0jirachi · x · 2026-10-03
Developer KyrieBlunders shared a hand-written Blackwell GEMM kernel using CuTe DSL on a B200, boosting throughput from 164 to 1401 TFLOP/s — roughly 97% of cuBLAS performance, showing custom kernels can nearly match vendor libraries.
More from Infra
- ERRC poster: entropy-reinvested residual correction for tensor-parallel LLM inference — PyTorch · 2026-10-03
- Dev platform cuts prices in half, making GitHub Action runners 12x cheaper — aniketmaurya · 2026-10-03
- Free 5% speedup: enable CustomAllReduce on SM120 to push tensor parallelism from TP=2 to TP=4 — TheZachMueller · 2026-10-03
- David Patterson: solar's near-vertical cost curve will power the singularity — davidpattersonx · 2026-10-03
- Local LLM setup: Strata lets a 7900XTX + 64GB RAM run Qwen Flash at 60 tok/s at 250K context — soyalemujica · 2026-10-03
- DGX Spark shortage derails $8,800 donation plan as buyer can't find stock — cyrus_zei · 2026-10-03