What breaks when AI agents run in production? Developer explores an SRE-style control layer

Fantastic-Sleep-3352 · reddit · 2026-09-05

A developer is exploring building an SRE-like control layer beneath AI agents that monitors trajectory and state, then decides whether to retry, replan, switch tools, downgrade models, restore state, escalate to a human, or stop — verifying task success before completion. He lists 12 validation questions for teams running agents in production, targeting pain points like runaway loops, token cost spikes, untraceable failures, missing replay/audit, and unclear human-approval thresholds, arguing this unglamorous reliability layer may matter more than another agent framework.

Related event: Developer seeks to build an SRE-style control layer for production AI agents(2 posts)→

Original post →

More from coding & agent

coding & agent channel →