What breaks when AI agents run in production? Developer explores an SRE-style control layer
Fantastic-Sleep-3352 · reddit · 2026-09-05
A developer is exploring building an SRE-like control layer beneath AI agents that monitors trajectory and state, then decides whether to retry, replan, switch tools, downgrade models, restore state, escalate to a human, or stop — verifying task success before completion. He lists 12 validation questions for teams running agents in production, targeting pain points like runaway loops, token cost spikes, untraceable failures, missing replay/audit, and unclear human-approval thresholds, arguing this unglamorous reliability layer may matter more than another agent framework.
More from coding & agent
- Dev builds Diablo-inspired browser ARPG Evergrow with Astra, open-sources playable prototype — Dimillian · 2026-09-06
- Astra vs Fable in practice: one just does it right, one has better design sense — weswinder · 2026-09-06
- The 4 layers of an agent system: failures are often architecture problems, not prompting problems — blaizedsouza · 2026-09-06
- OpenCUA author: the 450-step Paint demo we rejected 2 years ago is now doable by agents — andersonbcdefg · 2026-09-06
- One-line prompt: agent formats a messy Google Doc in 45 seconds — vaibhavbetter · 2026-09-06
- Codex Astra adds cross-context note persistence to fight context bloat in long tasks — dair_ai · 2026-09-06