Analyzing NVIDIA's NVFP4 Pre-training Approach

nrehiew_ · x · 2026-07-13

The post analyzes NVIDIA's **NVFP4 pre-training approach**, noting that it heavily borrows from previous Nemotron work: - **Hadamard transform**: Applied to weight gradient calculations to reduce the impact of outliers. - **Selective high precision**: Certain layers (like the final layer) maintain high precision because they require a larger dynamic range and mantissa than FP4. - **Stochastic rounding**: Uses stochastic rather than deterministic rounding in gradient calculations to prevent bias. To validate the approach, the team trained smaller models on up to 16T tokens, showing only about a 0.4% relative training loss gap compared to the BF16 baseline.

Related event: Deep Dive into NVIDIA's NVFP4 Quantization and Pretraining(2 posts)→

Original post →

More from Infra

Infra channel →