LinkedIn's GPU Pre-Ranker Fuses Graph Features for 50x Bigger Models at 120ms p99
_reachsumit · x · 2026-09-22
LinkedIn published Connected Content Retriever (arXiv:2609.22441), a GPU pre-ranking system for its Feed, where network-generated content accounts for over 70% of impressions and engagement. The pre-ranking layer must score tens of thousands of activities from a candidate index exceeding one billion within a 120 ms p99 latency budget.
Core design: a sorted-search GPU primitive joins dense graph affinity features (viewer-to-author) with document-level features stored on GPU at runtime in 5-10 ms, covering both first- and second-degree networks including 'stranger viral' content.
Impact: GPU-served scoring enabled a 50x scale-up of the ranking model's parameters and delivered a +2.5% lift while meeting the strict latency budget.
More from Infra
- 63% of Americans oppose data centers in their community, spanning both parties, pollster says — AlexTensor · 2026-09-22
- Toby Ord estimates $20M spent on AI's millennium prize result, $200M for solid data — tobyordoxford · 2026-09-22
- What engineers check before adding a new LLM provider — Rama_Surasani_ · 2026-09-22
- Dev reverse-engineers DLSS 5 neural rendering, reimplements it bit-exact in Vulkan at 7.8ms/1080p — bdsqlsz · 2026-09-22
- M5 Ultra Hits 3740 tok/s Prefill on Qwen, Nearly Double Overnight — EAccelerate_42 · 2026-09-22
- Running 2.78T-param Kimi K3 on a single CPU in 8 GB RAM, no framework — tom_doerr · 2026-09-22