Meta Scales Ads Recommendation Model to LLM Size, Doubling Training Efficiency
Meta_Engineers · x · 2026-08-05
Meta Engineering announced that its Generative Ads Recommendation Model (GEM), powering ads on Instagram and Facebook, is now training at LLM scale across thousands of latest-gen GPUs.
Because standard LLM infrastructure natively falls short for recommendation workloads, the team co-designed deep system optimizations:
- Custom Kernels: Developed tailored kernels like Jagged Flash Attention and applied ultra-low mixed precision, including MXFP8 attention.
- Parallelism: Built a topology-aware 5D parallelism architecture specifically mapped to their network hierarchy.
These engineering efforts yielded significant results: over 12 months, total training FLOPs scaled 4x while end-to-end training efficiency doubled, achieving a 20–25% Model FLOPs Utilization (MFU).
Related event: Meta Scales Ad Recommendation Models to LLM Size(3 posts)→
More from Infra
- AMD Claims Helios Rack Beats Nvidia's Vera Rubin by 30% in Tokens per Dollar — Beth_Kindig · 2026-08-05
- Burning $130K/Day? Unpacking DeepSeek API Token Volumes — teortaxesTex · 2026-08-05
- NSF Launches $100M Program for Regional AI Infrastructure Hubs — mkratsios47 · 2026-08-05
- Chutes AI Enforces TEE Verification: 8x RTX 5090s Beat Pro GPUs at 65% Lower Cost — markjeffrey · 2026-08-05
- engyai Launches Cheapest Kimi K3 API on OpenRouter, Cutting Costs by 50% — const_reborn · 2026-08-05
- LiquidAI's LFM2.5-2.6B Hits 82 tok/s Decode on Mac with 128K Context — helloiamleonie · 2026-08-05