How do you test production AI agents before real users touch them?

sixeyedhere · reddit · 2026-09-28

A developer opens a discussion on evaluating customer-facing AI agents beyond the usual "did the model answer well" LLM eval. Production agents raise harder questions: did the agent access correct data, call the right tool/API, follow authentication and permission rules, take the correct action, and know when to escalate? An example lending agent quoting ₹18,400 sounds right, but you must verify it matches the backend value, the customer was authenticated, the right account was fetched, and voice recognition didn't garble the amount. The author asks what teams actually use — manual QA, custom eval datasets, LLM-as-a-judge, unit tests for tools, observability platforms, production monitoring — and the biggest post-deployment failures they've seen.

Original post →

More from coding & agent

coding & agent channel →