Pretraining Large Language Models with NVFP4
NVIDIA, Felix Abecassis, Anjulie Agrusa, Dong Ahn, Jonah Alben, Stefania Alborghetti, Michael Andersch, Sivakumar Arayandi, Alexis Bjorlin, Aaron Blakeman, Evan Briones, Ian Buck, Bryan Catanzaro, Muya Chang, Jinhang Choi, Mike Chrzanowski, Eric Chung, Victor Cui, Steve Dai, Bita Darvish Rouhani, Carlo del Mundo, Deena Donia, Burc Eryilmaz, Henry Estela, Abhinav Goel, Oleg Goncharov, Yugi Guvvala, Robert Hesse, Russell Hewett, Herbert Hum, Ujval Kapasi, Brucek Khailany, Mikail Khona, Nick Knight, Alex Kondratenko, Ronny Krashinsky, Ben Lanir, Simon Layton, Michael Lightstone, Daniel Lo, Paulius Micikevicius, Asit Mishra, Tim Moon, Deepak Narayanan, Chao Ni, Abhijit Paithankar, Satish Pasumarthi, Ankit Patel, Mostofa Patwary, Ashwin Poojary, Gargi Prasad, Sweta Priyadarshi, Yigong Qin, Xiaowei Ren, Oleg Rybakov, Charbel Sakr, Sanjeev Satheesh, Stas Sergienko, Pasha Shamis, Kirthi Shankar, Nishant Sharma, Mohammad Shoeybi, Michael Siu, Misha Smelyanskiy, Darko Stosic, Dusan Stosic, Bor-Yiing Su, Frank Sun, Nima Tajbakhsh, Shelby Thomas, Przemek Tredak, Evgeny Tsykunov, Gandhi Vaithilingam, Aditya Vavre, Rangharajan Venkatesan, Roger Waleffe, Qiyu Wan, Hexin Wang, Mengdi Wang, Lizzie Wei, Hao Wu, Evan Wu, Keith Wyss, Ning Xu, Jinze Xue, Charlene Yang, Yujia Zhai, Ruoxi Zhang, Jingyang Zhu, Zhongbo Zhu
cs.CL, cs.AI, cs.LG
2025-09-30
NVIDIA pretrained a 12B hybrid Mamba-Transformer for 10T tokens in NVFP4; MMLU-Pro is 62.58% vs 62.62% for FP8, with relative loss error under 1% in the stable phase.
FP8 is already the default for large-scale pretraining. FP4 is the next cut: on Blackwell, Tensor Cores run FP4 math at about 2x (GB200) to 3x (GB300) the throughput of FP8, and operand memory is roughly halved. The catch is the tiny dynamic range. Over a long token horizon, quantization noise can accumulate and the run can drift off a higher-precision baseline.
No public 4-bit run had previously taken a billion-parameter LLM through a multi-trillion-token horizon. This report asks an algorithm question, not a wall-clock one: with NVIDIA's NVFP4 format and a small set of training tricks, can a 12B hybrid Mamba-Transformer track an FP8 baseline across 10T tokens.
NVFP4 changes three knobs relative to OCP MXFP4. Blocks shrink from 32 elements to 16. The block scale moves from a power-of-two UE8M0 factor to E4M3, which has a mantissa. A tensor-level FP32 scale sits on top. The block maximum can be recovered at near-FP8 fidelity; the rest of the block is stored as E2M1 (the discrete set 0, ±0.5, …, ±6). MXFP4's power-of-two scales can, in the worst case, waste the ±4 and ±6 bins and lose nearly a binade of range.
The format alone is not enough. Quantizing every linear layer to FP4 diverges. The recipe that holds at 12B / 10T is:
Embeddings, the output head, norms, nonlinearities, attention softmax, and the QK / AV batched GEMMs stay in the original precision. Master weights, gradient accumulation, and optimizer state stay FP32. Tensor-parallel reductions run in BF16.
The 12B model follows the Nemotron-H / Nemotron-Nano-12B-v2-Base hybrid layout (62 blocks: 6 self-attention, 28 FFN, 28 Mamba-2), sequence length 8192, batch 736, WSD schedule with a constant LR for 80% of training then a 20% decay, on 10T tokens. The control is FP8 pretraining on the same data. Downstream eval is always in BF16.
Relative loss error stays under 1% in the stable phase and widens to a little over 1.5% during decay. The slope change at 8T is the LR decay; the small jump at 9T is a data-blend switch. Task scores stay close:
| Task | FP8 | NVFP4 |
| MMLU-Pro 5-shot | 62.62 | 62.58 |
| MMLU | 77.36 | 76.57 |
| GSM8k CoT | 89.08 | 92.27 |
| MATH | 83.32 | 81.48 |
| HumanEval+ | 59.93 | 57.43 |
| MBPP+ | 59.11 | 55.91 |
The three coding numbers lag by about 2–3 points. The paper guesses MBPP+ is last-checkpoint noise and does not re-evaluate another checkpoint. To close the loss gap, switch to higher precision before decay: moving to BF16 after 8.2T tokens (about 18% of training) matches FP8; switching only in the last <1% of steps still cuts relative error from 1.5% to about 0.5%.
On an 8B model trained for 1T tokens, NVFP4's relative loss error versus BF16 is about 1.5%, MXFP4 about 2.5%. MXFP4 needs 36% more tokens (1.36T vs 1T) to match NVFP4's loss.
Dropping any one of the four ingredients hurts 12B / 10T. Smaller models and shorter horizons often do not need the full set.
This is the longest publicly documented 4-bit pretraining run. The algorithm holds to 10T, and Transformer Engine already landed the kernels. For teams training on Blackwell, GEMM-heavy linear layers can drop to FP4 while attention and the last blocks stay wider.
It is still mixed precision. Sixteen percent of linear layers remain BF16, and the attention path is untouched. The report is explicit that it studies numerical feasibility, not end-to-end step time. Real speedups will depend on how much of the step is linear GEMMs versus communication.
Against MXFP4, NVFP4 converges better at a fixed token budget. The 36% extra tokens MXFP4 needs is a number a training budget can use.
The coding gap is not cleanly explained, only waved at as last-checkpoint noise. All downstream numbers are BF16 eval, so FP4 inference quality is unknown.
The recipe is scale-sensitive: 12B / 10T needs all four pieces, 1.2B often does not. MoE, longer horizons, and quantizing attention are untested. The fact that the last layers must stay wide means FP4 still cannot cover the dynamic range near the output.
Hadamard is restricted to Wgrad to keep 2D weight scaling consistent. Whether larger models will also need it on Fprop and Dgrad is open. The tensor-level FP32 scale costs an extra full-tensor amax pass; that is not free in a real stack.