How do you test production AI agents before real users touch them?
sixeyedhere · reddit · 2026-09-28
A developer opens a discussion on evaluating customer-facing AI agents beyond the usual "did the model answer well" LLM eval. Production agents raise harder questions: did the agent access correct data, call the right tool/API, follow authentication and permission rules, take the correct action, and know when to escalate? An example lending agent quoting ₹18,400 sounds right, but you must verify it matches the backend value, the customer was authenticated, the right account was fetched, and voice recognition didn't garble the amount. The author asks what teams actually use — manual QA, custom eval datasets, LLM-as-a-judge, unit tests for tools, observability platforms, production monitoring — and the biggest post-deployment failures they've seen.
More from coding & agent
- GPT Researcher drops embeddings: LLM-based retrieval lifts relevant context 59% at same cost — hwchase17 · 2026-09-28
- DHH: Programmers who deny AI's paradigm shift are the ones truly at risk — CSProfKGD · 2026-09-28
- EmDash 1.0 ships: open-source Astro CMS with MCP server and decentralized plugin registry — irvinebroque · 2026-09-28
- Merge ships five new connectors for Agent Handler, including Power BI and Dynamics 365 — shensi · 2026-09-28
- Developer builds AI customer support agent that remembers past conversations with Hindsight — Own_Train718 · 2026-09-28
- A PR arrived saying "I wouldn't expect humans to understand this code" — smlpth · 2026-09-28