NVFP4: Inside the Training Tricks That Let NVIDIA Pretrain LLMs in FP4
bycloud · youtube · 2026-09-22
A new video from bycloud breaks down NVFP4, the low-precision format NVIDIA is betting on for its upcoming Vera Rubin GPUs. Key points:
- FP4 training is far harder than just using fewer bits; NVIDIA relies on tricks like stochastic rounding and 2D block scaling to keep pretraining stable.
- The analysis draws on primary sources: the NVFP4 pretraining paper (arXiv:2509.25149), NVIDIA's Quantization-Aware Distillation report for recovering FP4 inference accuracy, the Nemotron 3 white paper, and the Nemotron 3 Ultra MoE hybrid Mamba-Transformer technical report.
- Low precision is NVIDIA's core lever for cutting training and inference cost, and it is now shaping GPU architecture itself, with the DeepSeek-V4 efficiency paper cited as further evidence of the industry-wide push toward low-precision training.
More from Infra
- UK's most powerful government AI supercomputer cost £225m — same as one bridge — emax · 2026-09-22
- exe.dev wins over developers: SSH into root VMs, plus the underrated Shelley coding agent — davidcrawshaw · 2026-09-22
- NVIDIA hosts trilateral meeting as Artificial Analysis becomes Korea sovereign AI evaluator — ArtificialAnlys · 2026-09-22
- Lisa Su's Generational Run: AMD Market Cap Up 300x to $1T, Data Center CPU Share Hits 40%+ — xiaosun86 · 2026-09-22
- Claude Code's cross-session messaging is just a JSON directory and one Unix socket per session — LeonKohli · 2026-09-22
- SB Energy's $50B-Valuation IPO for World's Largest Data Center Project Delayed on Weak Investor Demand — firstadopter · 2026-09-22