Meta Scales Ads Recommendation Model to LLM Size, Doubling Training Efficiency
Meta_Engineers · x · 2026-08-05
Meta's engineering team details the latest upgrades to their Generative Ads Recommendation Model (GEM), the engine powering ads across Instagram and Facebook. The model is now training at LLM scale across thousands of next-gen GPUs.
Because standard LLM infrastructure doesn't natively fit recommendation workloads, Meta engineered custom solutions:
- Custom Kernels: Co-designed Jagged Flash Attention.
- Mixed Precision: Utilized ultra-low MXFP8 precision.
- Parallelism: Built a topology-aware 5D parallelism architecture tailored to their network hierarchy.
These optimizations allowed Meta to scale total training FLOPs 4x in 12 months while doubling end-to-end training efficiency to 20–25% MFU.
Related event: Meta Scales Ad Recommendation Models to LLM Size(3 posts)→
More from Infra
- YC-Backed Lamb Labs Claims 63x Higher Efficiency for AI Inference Chips vs GPUs — ycombinator · 2026-08-05
- DSpark Open-Sources Speculative Decoding Path for Kimi K3 — ying11231 · 2026-08-05
- US Drafts Ban on Chinese Datacenter Components as Europe Pushes for Tech Sovereignty — nordicinst · 2026-08-05
- Cursor Releases Mixture-of-Kittens Megakernel for MoE, Claims Nearly 2x TFLOP/s — CapnHat · 2026-08-05
- NVIDIA Tutorial: Building Fully Local Autonomous Agents on Jetson — NVIDIA Developer · 2026-08-05
- YC Launches Caution Hosting: Secure Enclaves for Sensitive AI Workloads — ycombinator · 2026-08-05