Solid evals, still blind: how Conviva caught agent failures users quietly gave up on

Automatic-Mirror7324 · reddit · 2026-09-10

Conviva's analytics agent served 20 major streaming providers during the 2026 World Cup with tens of millions of concurrent users. Their eval suite scored well but caught almost nothing — because test cases only came from loud failures, while the expensive ones were silent: a conversation scoring 0.94+ on helpfulness while the agent quietly omitted an optional search parameter and the user left to do it manually.

What worked instead:

Original post →

More from coding & agent

coding & agent channel →