New B300 FP4 Attention Kernels Deliver 1.69x Speedup

tuananh_org · reddit · 2026-07-14

Someone shared a set of new FP4 attention kernels for the B300, claiming up to a 1.69x speed improvement compared to FA4.

This update is fundamentally about inference/operator-level performance optimization rather than the models themselves. If accurate, it indicates substantial room for optimizing attention implementations on new hardware platforms, particularly along low-precision pathways.

Original post →

More from Infra

Infra channel →