Beating cuBLAS by 4.7%: NVFP4 Kernels Hand-Built on GB300, 100% Claude-Generated

pranjalssh · x · 2026-08-16

Two years ago, the author entered the GPU world with a blog on outperforming cuBLAS on the H100 — a post that unexpectedly sparked a mini wave of hand-written Blackwell kernels: others reached 99% of cuBLAS with MXFP8, beat it with BF16, and the Modular team did it too. Now he's back with a sequel: an NVFP4 GEMM built from scratch on a GB300 node (CUDA 13.1) in plain CUDA plus inline PTX, with no framework dependency.

He hopes this motivates people to own their kernels as ordinary code rather than treating them as black boxes.

Original post →

More from Infra

Infra channel →