SREGym benchmark launches: GPT-5.6 leads agents at fixing real production SRE failures
tianyin_xu · x · 2026-09-09
Researchers at UIUC launched SREGym, a benchmark that tests whether AI agents can resolve real-world SRE issues in live system environments, covering metastable failures, misconfigurations, and more. It will be presented at the Frontier Data Summit in SF on Oct 8, alongside teams behind OSWorld 2.0, Terminal Bench, T² Scaling Laws, and other benchmarks.
Top results on the curated 21-fault cohort:
- Codex + GPT-5.6 Sol (max): 95.2% diagnosis, 85.7% mitigation, 81.0% end-to-end, averaging 1.42M tokens per run
- Claude Code + Claude Opus 5: 92.1% diagnosis, 82.5% mitigation, 76.2% E2E
- Codex + GPT-5.6 Terra (max): 69.8% E2E
- Codex + GPT-5.6 Luna (max): 68.3% E2E but a hefty 2.74M tokens per run
Metrics include diagnosis and mitigation success rates, end-to-end accuracy, and time-to-diagnose/mitigate. The project is open source and accepting community submissions.
More from coding & agent
- Plausible Analytics MCP server lets AI assistants query website traffic stats — modelcontextprotocol · 2026-09-09
- Human Design MCP connector brings bodygraph chart analysis to AI assistants — modelcontextprotocol · 2026-09-09
- One change cut Claude Code tokens 3x: 10.4M→3.7M, $9.21→$2.81 with InsForge — Roger_M_Taylor · 2026-09-09
- Google releases 1-hour graph engineering tutorial: from single agent to self-improving 24/7 systems — Roger_M_Taylor · 2026-09-09
- Personal AI Workflows Are the New Company Assets: Workers May Leave With Their AI Agents — aigclink · 2026-09-09
- Agent given a computer builds its own simulation-with-agents, in a simulation — Confident_Salt_8108 · 2026-09-09