1 in 4 AI coach chats failed with zero alerts — Amplitude opens up Agent Analytics to close the agent QA blind spot
aakashgupta · x · 2026-09-15
YC company HYBRID's AI coach botched 1 in 4 chats while logging zero errors: a misfiring tool retry still produced answers, so analytics saw complete sessions and observability saw clean traces — only a conversion drop could have exposed it, far too late.
The author argues this is the blind spot of most agent teams: analytics shows sessions happened, observability shows traces completed, but neither answers whether replies were good, trusted, or drove conversion. Amplitude hit the same gap in 2025, built Agent Analytics internally, and is now opening it up:
- Every trace scored for task success, failure type, and tool-call errors
- Every conversation becomes a measurable conversion event — rank model providers on conversion, cluster by intent, fix the costliest failing topic first
Three customer results: The Economist runs its Lens assistant at a 96.9% task success rate with weekly failures down 84%; HYBRID athletes using the fixed AI coach retain 4x more; Cashbook found a bad first AI answer quietly cost 15% retention across its app. Takeaway: reply quality is becoming a deliberately tracked metric, and teams counting it now catch failures before the revenue chart does.
More from coding & agent
- Hackathon Project Skillio Turns Vision Pro Into an AR Screw-Driving Coach Agent — seanmcdonaldxyz · 2026-09-15
- Let the Model Propose, the Flow Validate, a Human Approve: Drafting Agent Checkpoints — WirelessLife · 2026-09-15
- Devin-powered daily briefs: how one dev uses AI agents to automate information discovery — bendee983 · 2026-09-15
- Claude writes 80% of code at Anthropic, CI jobs up 25x in six months — addyosmani · 2026-09-15
- Voice agent production postmortem: 7 behavioral bugs found only in real call logs — authentic_developer · 2026-09-15
- Dev Sells .md Domains for Two Unbuilt AI Ideas: codebase.md ($60) and discover.md ($500) — iannuttall · 2026-09-15