Meta paper: only 50-60% of recommendation training time actually trained before optimizations
_reachsumit · x · 2026-10-02
A Meta team has published an arXiv paper on fleet-scale inefficiency in recommendation system training, where their largest workloads process tens of billions of examples daily on thousands of GPUs. Before the work, only 50-60% of end-to-end wall time advanced training on new data.
Key contributions:
- Effective Training Time (ETT%): an operational framework to instrument lost time, attribute it to independently owned infrastructure components, and expose work repeated across job restarts
Optimizations span the full training stack: communication elimination and pipeline overlap during trainer initialization, dynamic-shape handling and autotuning pruning, reusable PyTorch 2 compilation caches, asynchronous checkpointing, standalone model publishing, and reduced recovery cost. The team evaluated each optimization on representative models and measured fleet-wide ETT% improvements.
More from Infra
- awesome-jev indexes 700 production tools around TypeSafe AI's decision model Jev — Remarkable-Gur719 · 2026-10-02
- smolvm v1.22 ships near-instant VM resume for undoing agent actions, 6.5k stars — LoganGrasby · 2026-10-02
- Broadcom to lend Anthropic up to $42B for AI chips, eyeing top customer slot by 2027 — rohanpaul_ai · 2026-10-02
- Dev open-sources GPT-2-tools to run original 1.5B GPT-2 XL locally on CPU — MikePFrank · 2026-10-02
- CoreWeave launches serverless GPUs: hourly-billed, no contract, private preview — altryne · 2026-10-02
- Japan plans $140B AI data center push; Huawei ships Mate90 with Taou chip — 创业邦 · 2026-10-02