1 in 4 AI coach chats failed with zero alerts — Amplitude opens up Agent Analytics to close the agent QA blind spot

aakashgupta · x · 2026-09-15

YC company HYBRID's AI coach botched 1 in 4 chats while logging zero errors: a misfiring tool retry still produced answers, so analytics saw complete sessions and observability saw clean traces — only a conversion drop could have exposed it, far too late.

The author argues this is the blind spot of most agent teams: analytics shows sessions happened, observability shows traces completed, but neither answers whether replies were good, trusted, or drove conversion. Amplitude hit the same gap in 2025, built Agent Analytics internally, and is now opening it up:

Three customer results: The Economist runs its Lens assistant at a 96.9% task success rate with weekly failures down 84%; HYBRID athletes using the fixed AI coach retain 4x more; Cashbook found a bad first AI answer quietly cost 15% retention across its app. Takeaway: reply quality is becoming a deliberately tracked metric, and teams counting it now catch failures before the revenue chart does.

Original post →

More from coding & agent

coding & agent channel →