NVIDIA to Demo NVFP4 Pretraining Recipes Approaching BF16 Quality
PyTorch · x · 2026-09-01
NVIDIA researchers will present the latest NVFP4 pretraining recipes at PyTorch Conference North America 2026, aiming to accelerate large-scale LLM training while closing the quality gap to BF16. The October 21 session will cover recipe design, kernel choices, and the API surface required for native PyTorch integration. Dense linear NVFP4 training is currently available via TorchAO, with the recipes being upstreamed into the PyTorch ecosystem through TorchAO and TorchTitan.
More from Infra
- TensorSharp vs llama.cpp: Qwen 3.8 Flash Next Benchmarks — fuzhongkai · 2026-09-01
- 2 engineers + AI designed a working LLM chip in 2 weeks, no human in the loop — 新智元 · 2026-09-01
- Why did increasing context size increase speed in Llama.cpp? — satnl · 2026-09-01
- AI inference demand surges again, supply brutally outpaced by token growth — Baconbrix · 2026-09-01
- Warp founder predicts cloud-based collaborative factories for all companies within a year — charlieholtz · 2026-09-01
- JPM: 1GW of AI Infrastructure Costs $40-45B, Frontier Labs Make ~$30B per GW — zephyr_z9 · 2026-09-01