Dev seeks blueprint for agent that triages incidents across PagerDuty, Datadog, GitLab and Slack
cruelcaricature · reddit · 2026-09-14
A Reddit developer is asking for architecture advice on building an agentic system for production incident triage: when a Jenkins failure or PagerDuty alert surfaces in Slack, the agent should inspect the relevant Jenkins job and Datadog monitors, infer which repo caused the issue, fix it, and open an MR.
He already uses RAG over Confluence docs and Slack conversations, but is stuck on how to structure multi-repo code and dependency information for agent querying, and asks for open-source tools that map interrelated repo dependencies.
More from coding & agent
- Anthropic demos Claude as on-call engineer: 15-minute incident triage in Slack — xiaohu · 2026-09-14
- Agent hacks stem from scraping local private files, not alignment failure, researcher argues — ryunuck · 2026-09-14
- Diversity-Aware Skill Routing Uses DPP to Cut Redundant LLM Agent Skill Picks — Wang Wei · 2026-09-14
- Contextual Bandit Algorithms Route Prompts to LLM Experts with Sublinear Regret — Wang Wei · 2026-09-14
- Step-by-Step Guide to Becoming an Inference Optimization Engineer, From Quantization to Speculative Decoding — ashishllm · 2026-09-14
- DeepDeck ships a WebMCP directory with pinned revisions: build inspectable browser tools for sites without APIs — j032 · 2026-09-14