OpenAI reports 4.1x training throughput and 95%+ cluster utilization gains

LiamFedus · x · 2026-09-16

OpenAI's Liam Fedus shared the lab's infrastructure goal of "science per GPU-hour." On real workloads they achieved 4.1× training throughput vs. their Megatron baseline, 2.5× faster decoding, and 95%+ cluster utilization, contributing improvements back to Megatron-LM, SGLang, and Miles.

A follow-up highlighted the scientific payoff: lab data + midtraining + RL raised X-ray diffraction analysis success from 2.7% to 55.3% (20×) on 134 difficult samples, scored by model judges calibrated against human experts, with strong scaling as RL is added.

Original post →

More from Infra

Infra channel →