mamf-finder adds FP8/MXFP4/NVFP4 support for real GPU TFLOPS benchmarking
StasBekman · x · 2026-10-03
Stas Bekman updated mamf-finder in the ml-engineering repo to support benchmarking peak matrix-multiply throughput across formats: bfloat16, float16, float32, float8e4m3fn, float8e4m3fnuz, mxfp8, mxfp4, and nvfp4 — letting you measure your real 100% TFLOPS for almost any workload.
The accompanying docs explain a subtle measurement pitfall: kernel time is not a function of tensor shapes alone. Input values affect how hard the chip works — high-entropy bit patterns flip more transistors than sparse or all-zero patterns, raising power draw, and once it approaches the power limit, clocks drop. Hence locking GPU/memory clocks matters, as mamf-finder does.
More from Infra
- Supabase Launches Compute: Long-Running Servers for Heavy Workloads — dshukertjr · 2026-10-03
- Underdog raises from a16z to put private AI on devices you already own — Thom_Wolf · 2026-10-03
- Details emerge on DeepSeek's custom chip: skip-scale optimization cuts power, boosts frequency — teortaxesTex · 2026-10-03
- SkyPilot raises $20M seed to unify GPU scheduling across clouds and clusters — skypilot_org · 2026-10-03
- The Epochalypse: 32-bit systems will think it's 1901 after Jan 19, 2038, 03:14:07 UTC — iannuttall · 2026-10-03
- Dev builds free 3D interactive guide to explain how unified GPU/CPU systems power AI data centers — ghumare64 · 2026-10-03