Building an SRE-Style Control Layer for Production AI Agents, Seeking Validation
Fantastic-Sleep-3352 · reddit · 2026-09-06
A developer proposes building an SRE-like control layer underneath production AI agents — not another framework — that watches agent trajectories, diagnoses failures, and decides whether to retry, replan, switch tools, downgrade models to cut cost, restore state, escalate to humans, or stop, verifying task completion along the way. He lists common pain points: runaway tool calls, loops, retry storms, token cost explosions, and unrecoverable long-running failures, and asks 12 validation questions of developers running agents in production about tooling (LangSmith, Langfuse, Arize), replayability, and autonomy thresholds.
More from coding & agent
- Self-building IDE bb gains agent chat rooms; dev says don't roll your own orchestration — andpoul · 2026-09-06
- Giving multiple agents a shared chat room to pool info and coordinate — andpoul · 2026-09-06
- Kornia's DoG/Hessian now match OpenCV quality, with Astra improving results in one shot — ducha_aiki · 2026-09-06
- Coddle launches an agentic workspace that covers the full product development lifecycle — saheedniyi_02 · 2026-09-06
- 50+ autonomous bots build a 215-parcel virtual world with building codes in one day — Daniel_Farinax · 2026-09-06
- Tip: pair Codex with Muse Spark to cut token burn, free on opencode — BLUECOW009 · 2026-09-06