Analyzing NVIDIA's NVFP4 Pre-training Approach
nrehiew_ · x · 2026-07-13
The post analyzes NVIDIA's **NVFP4 pre-training approach**, noting that it heavily borrows from previous Nemotron work: - **Hadamard transform**: Applied to weight gradient calculations to reduce the impact of outliers. - **Selective high precision**: Certain layers (like the final layer) maintain high precision because they require a larger dynamic range and mantissa than FP4. - **Stochastic rounding**: Uses stochastic rather than deterministic rounding in gradient calculations to prevent bias. To validate the approach, the team trained smaller models on up to 16T tokens, showing only about a 0.4% relative training loss gap compared to the BF16 baseline.
Related event: Deep Dive into NVIDIA's NVFP4 Quantization and Pretraining(2 posts)→
More from Infra
- Kimi K3 costs $4.65 per run and delivers 2.8× more work per dollar than Fable 5 — FinanceYF5 · 2026-07-21
- UK AI datacentres face backlash over heat, noise and land use — nordicinst · 2026-07-21
- Fluidstack raises $830M at $7.5B valuation as Anthropic backs a $50B compute buildout — rohanpaul_ai · 2026-07-21
- Early Krea2 Gradio WebUI targets 6GB low-VRAM local runs — Fluid_Kaleidoscope17 · 2026-07-21
- Z.AI starts running a 1GW AI data center built entirely on domestic chips — Polymarket · 2026-07-21
- Local models feel far more capable once paired with the right harness — Soft-Barracuda8655 · 2026-07-21