PyTorch and AMD Optimize FP8 Training: Over 13% Throughput Gain
PyTorch · x · 2026-08-14
A recent PyTorch blog post details how AMD has upstreamed FP8 training optimizations into PyTorch/TorchTitan and PyTorch/TorchAO, enabling out-of-the-box FP8 training on AMD Instinct GPUs.
Key optimization results include:
- Dense Models: Llama3-8B achieves a 13.4% throughput gain over BF16.
- MoE Architectures: For models like DeepSeek-V3 671B, fusing Triton quantization kernels recovered 89% of the FP8 quantization overhead, with individual kernel optimizations delivering up to a 6.2x speedup.
These improvements were achieved by systematically fusing Triton quantization kernels, enabling grouped GEMM for ROCm, and building a Triton fusion pipeline.
More from Infra
- The New Frontier of CS: Insane Potential of Shared Memory Across Computers — lauriewired · 2026-08-14
- Tinkering with Mining Cards: 50% Speed Boost for llama.cpp on CMP 170HX — fragment_me · 2026-08-14
- OpenAI Previews Ultrafast Mode: GPT-5.6 Sol Hits 14x Speeds — OpenAI · 2026-08-14
- Detecting Performance Regressions Using ML and Hardware Counters — ZeroDark_Hereford · 2026-08-14
- Polymarket Prices Nvidia at 73% Chance to Be World's Largest Company by 2026 — Polymarket · 2026-08-14
- Debunking Seven Common Myths About AI Data Centers' Environmental and Economic Impact — Dan_Jeffries1 · 2026-08-14