mamf-finder adds FP8/MXFP4/NVFP4 support for real GPU TFLOPS benchmarking

StasBekman · x · 2026-10-03

Stas Bekman updated mamf-finder in the ml-engineering repo to support benchmarking peak matrix-multiply throughput across formats: bfloat16, float16, float32, float8e4m3fn, float8e4m3fnuz, mxfp8, mxfp4, and nvfp4 — letting you measure your real 100% TFLOPS for almost any workload.

The accompanying docs explain a subtle measurement pitfall: kernel time is not a function of tensor shapes alone. Input values affect how hard the chip works — high-entropy bit patterns flip more transistors than sparse or all-zero patterns, raising power draw, and once it approaches the power limit, clocks drop. Hence locking GPU/memory clocks matters, as mamf-finder does.

Original post →

More from Infra

Infra channel →