Meta paper: only 50-60% of recommendation training time actually trained before optimizations

_reachsumit · x · 2026-10-02

A Meta team has published an arXiv paper on fleet-scale inefficiency in recommendation system training, where their largest workloads process tens of billions of examples daily on thousands of GPUs. Before the work, only 50-60% of end-to-end wall time advanced training on new data.

Key contributions:

Optimizations span the full training stack: communication elimination and pipeline overlap during trainer initialization, dynamic-shape handling and autotuning pruning, reusable PyTorch 2 compilation caches, asynchronous checkpointing, standalone model publishing, and reduced recovery cost. The team evaluated each optimization on representative models and measured fleet-wide ETT% improvements.

Original post →

More from Infra

Infra channel →