OpenAI reports 4.1x training throughput and 95%+ cluster utilization gains
LiamFedus · x · 2026-09-16
OpenAI's Liam Fedus shared the lab's infrastructure goal of "science per GPU-hour." On real workloads they achieved 4.1× training throughput vs. their Megatron baseline, 2.5× faster decoding, and 95%+ cluster utilization, contributing improvements back to Megatron-LM, SGLang, and Miles.
A follow-up highlighted the scientific payoff: lab data + midtraining + RL raised X-ray diffraction analysis success from 2.7% to 55.3% (20×) on 134 difficult samples, scored by model judges calibrated against human experts, with strong scaling as RL is added.
More from Infra
- MLPerf Inference v6.1 results imminent: 30 submitters, new accelerators and platforms — TheKanter · 2026-09-16
- Lambda runs 19 nodes in a 16-node power budget with NVIDIA DSX, +24% throughput — TheZachMueller · 2026-09-16
- NVIDIA unveils DSX AI Factory platform to maximize output per megawatt — nvidia · 2026-09-16
- Claim: run a 125B MoE at 22 tok/s on $320 of used GPUs with llama.cpp — cephaloform · 2026-09-16
- Astera Labs' Leo X-Series memory controllers boost long-context inference TTFT by up to 62% — BenBajarin · 2026-09-16
- Dev measures under 4.5GB VRAM peak throughout full local generation — cocktailpeanut · 2026-09-16