Meta Scales Ads Recommendation to LLM Level, Doubling Training Efficiency
Meta_Engineers · x · 2026-08-05
Meta officially shared that its Generative Ads Recommendation Model (GEM), powering ads on Instagram and Facebook, is now training at LLM scale across thousands of latest-gen GPUs. Because standard LLM infrastructure doesn't natively fit recommendation workloads, Meta co-designed several custom low-level components:
- Custom Kernels: Including Jagged Flash Attention (JFA), BlockAttention, and ultra-low mixed precision with MXFP8.
- Architecture: A topology-aware 5D parallelism architecture tailored to their network.
Through these optimizations, Meta scaled total training FLOPs 4x in 12 months while doubling end-to-end training efficiency to 20–25% MFU.
Related event: Meta Scales Ad Recommendation Models to LLM Size(3 posts)→
More from Infra
- Microsoft Guides Azure Growth to Accelerate to 45% in FQ1 — Beth_Kindig · 2026-08-05
- AMD Claims Helios Rack Beats Nvidia's Vera Rubin by 30% in Tokens per Dollar — Beth_Kindig · 2026-08-05
- Burning $130K/Day? Unpacking DeepSeek API Token Volumes — teortaxesTex · 2026-08-05
- NSF Launches $100M Program for Regional AI Infrastructure Hubs — mkratsios47 · 2026-08-05
- Chutes AI Enforces TEE Verification: 8x RTX 5090s Beat Pro GPUs at 65% Lower Cost — markjeffrey · 2026-08-05
- engyai Launches Cheapest Kimi K3 API on OpenRouter, Cutting Costs by 50% — const_reborn · 2026-08-05