Strangely, GPU matmuls run faster on 'predictable' data: Horace He explains

goyal__pramod · x · 2026-09-25

Horace He recounts a counterintuitive finding: CUTLASS's profiler showed 288 TFLOPS vs CuBLAS's 258 on an 8192³ matmul, but the gain vanished in Python. Ablations revealed the profiler initializes inputs with integers only — zeros hit 295 TFLOPS while randn inputs drop to 257. The values' distribution directly affects matmul runtime, and the post explores why from first principles.

Original post →

More from Infra

Infra channel →