AMD to present MXFP8 pretraining scaling on 1K+ MI355X GPUs at PyTorchCon 2026
PyTorch · x · 2026-09-25
PyTorchCon North America 2026 preview: AMD's Liz Li and Shekhar Pandey will present scaling MXFP8 pretraining on 1K+ AMD Instinct MI355X GPUs, covering MXFP8 operator optimization in TorchAO, integration with TorchTitan for an end-to-end upstream LLM training stack, plus Triton and FlyDSL implementations and lessons from numerical validation, convergence and performance tuning. The conference runs Oct 20-21 in San Jose.
More from Infra
- VeriTile embeds Triton GPU kernels in Lean, with AI agents writing machine-checked correctness proofs — KaiyuYang4 · 2026-09-25
- Merge Gateway Launches Batch Inference at ~50% of Standard Prices — shensi · 2026-09-25
- New deep-dive article on scaling LLM inference in production — abhijithneil · 2026-09-25
- Lambda engineer shares local inference build rule: 27B models need 24-32GB VRAM — TheZachMueller · 2026-09-25
- Pokee AI demos 36B agent model running fully local on Snapdragon X2 Elite with 32GB RAM — Kyrannio · 2026-09-25
- AI energy startup Parallax launches with $117m from Founders Fund, Lux, Greylock and others — graceisford · 2026-09-25