Meta Speeds Up with Normalization Fusion

PyTorch · x · 2026-07-11

The post shares Meta's optimization strategies for normalization fusion, aiming to alleviate memory bottlenecks caused by **Normalization layers** in large models and recommendation systems. The core approach fuses normalization operators directly with **GEMM** and **Attention kernels**. By running normalization calculations on CUDA Cores in parallel with the Tensor Core execution pipeline, it reduces stalling caused by tiling differences. Mentioned solutions include **Lazy Pre-Norm**, **Multi-CTA Norm Fusion**, and **FlashNormAttention**, reportedly delivering significant speedups on real recommendation traffic and the **NVIDIA B200**.

Original post →

More from Infra

Infra channel →