Optimizing NVFP4 Blockscaled GEMM on RTX Pro 6000 Blackwell

HanGuo97 · x · 2026-08-11

Colfax Research published a technical article on optimizing NVFP4 blockscaled GEMM for the NVIDIA RTX Pro 6000 Blackwell GPU (SM120).

The article details optimization strategies for small and large problem shapes (such as wave quantization and L2 cache thrashing), along with a series of micro-optimizations. It ultimately achieves compute throughput gains of 4% to 40% across various matrix sizes, peaking at 1666 TFLOP/s for 16k shapes with an 83% utilization rate.

Original post →

More from Infra

Infra channel →