Meta's hybrid GPU-CPU retrieval serves personalization over a billion docs plus 20x CPU inventory
_reachsumit · x · 2026-09-21
Meta details a production-deployed hybrid GPU-CPU co-serving system for ultra-large-scale personalized search. A deep GPU pathway fuses retrieval and interaction pre-ranking over a billion-document pool, while a breadth-first CPU pathway searches an inventory 20x larger with lightweight personalized scoring. A full-system A/B test against the legacy CPU-only setup improved both model-scored relevance and substantive engagement.
More from Infra
- Personal AI Agents Are the Biggest Driver of the Sudden NAND Demand Surge — zephyr_z9 · 2026-09-21
- How Grammarly's Superhuman Serves 100B+ LLM Requests a Week — jefrankle · 2026-09-21
- Local LLM VRAM sweet spot: is a single 32GB R9700 better than adding a second card? — endgamedos · 2026-09-21
- Jensen Huang says data center water use is a myth: new cooling systems evaporate less than a pool — rohanpaul_ai · 2026-09-21
- New Paper Achieves Sub-Second Interactive Diffusion on Consumer GPUs — chaumian · 2026-09-21
- Jev Sparse Attention Cuts MiniMax H3 Video Generation Time by 40% — Repulsive_Gap_1678 · 2026-09-21