Lyft's Real-World User Evaluation Loop

AI Engineer · youtube · 2026-07-19

Nick Ung discusses how Lyft builds closer-to-production evaluations for its customer service and agent systems.

The core issue is that many offline evals are too "easy"—the "customers" in tests are just standard LLMs that completely lack the frustration, tangents, and adversarial nature of real users, leading to unanticipated failure modes post-launch.

Lyft's approach involves:

Their customer service agent handles roughly a third of Lyft's support queries, scaling to millions of conversations per month. This case study highlights how to build evaluation as a sustainable engineering system, rather than focusing on a single model.

Original post →

More from coding & agent

coding & agent channel →