Uber Eats ranking models serve 8M predictions/sec: how Uber scales ML feature consistency
AxSaucedo · x · 2026-09-29
Uber engineering details its feature logging framework for tackling train-serving skew, where features seen in production (e.g. en vs en-US, differing computation logic) diverge from training data, hurting model quality and reliability.
Key points:
- Uber Eats ranking models serve 8 million predictions per second, making feature freshness and consistency critical
- Online features span user session data, restaurant metadata, and precomputed behavioral features (past clicks, orders)
- The goal: make inference-time features the single source of truth for the next training iteration, building a training-data flywheel directly from inference pipelines
- Inconsistencies reduce model effectiveness, increase debugging time, and can impact service reliability
More from Infra
- antirez: give companies a fleet of DGX Sparks and Mac Ultras to survive API quota limits — antirez · 2026-09-29
- Ornith-1.5 releases DFlash draft checkpoints for 9B/397B/35B-A3B speculative decoding — jacek2023 · 2026-09-29
- Jensen Huang says Nvidia stock is undervalued, plans hundreds of billions in buybacks — firstadopter · 2026-09-29
- Hugging Face Transformers now runs llama.cpp GGUF quants natively on your laptop — ariG23498 · 2026-09-29
- Qdrant unveils Constella research preview: swap query embedding models without re-embedding your docs — qdrant_engine · 2026-09-29
- 124M model with a 65B embedding sparks the AFED disaggregation joke — YouJiacheng · 2026-09-29