Ops Control Plane for Agent Clusters

C00LDude6ix9ine · reddit · 2026-07-19

Cartha has built an **SDK-first control plane** for AI agent clusters, aiming to solve the "hard to operate when agents scale" problem. It provides: - **Complete tracing**: Records memory, tool calls, LLM steps, policies, and costs per run, supporting parent-child multi-agent chains and run comparisons. - **Strictly isolated scoped memory**: Controls visibility by user / agent / team / org, rather than just "storing everything". - **Actionable cost analysis**: Calculates costs by agent, tool call, and completed task, factoring in retries and failure rollbacks. - **Governance and replay**: Includes a policy gate, historical action replay, and failure analysis; it is MCP/A2A friendly and framework-agnostic. The author is explicitly looking for real-world agent users to poke holes in it, focusing on validating DX, whether the abstractions fit multi-agent loops, and if any critical circuit breakers are missing.

Related event: Cartha Launches Ops Control Plane for Agent Fleets(2 posts)→

Original post →

More from coding & agent

coding & agent channel →