SREGym from UIUC benchmarks SRE agents on real outages: GPT-5.6 Sol leads at 81% E2E
tianyin_xu · x · 2026-09-11
- SREGym is a new benchmark from UIUC that drops AI agents into live system environments with real-world SRE problems, including metastable failures and misconfigurations; submissions run via GitHub, with a launch at Snorkel AI's Frontier Data Summit (Oct 8, SF).
- Leaderboard on the curated 21-fault SREGym-Lite cohort: Codex + GPT-5.6 Sol (max) tops the board at 95.2% diagnosis / 85.7% mitigation / 81.0% end-to-end, with TTD 211s, TTM 397s and 1.42M tokens per run.
- Claude Code + Claude Opus 5 ranks second at 76.2% E2E; GPT-5.6 Terra (max) and Luna (max) follow at 69.8% and 68.3%, with Luna burning 2.74M tokens per run. A medium-effort Sol config hits 58.7% E2E at just 0.77M tokens.
- Takeaway: even the best agent fully resolves only 81% of realistic production faults, and accuracy-vs-token-cost tradeoffs are nearly 2x between configs. The project ships a dataset registry for evaluating agents on standard third-party benchmarks.
Related event: UIUC releases SREGym benchmark, GPT-5.6 leads(2 posts)→
More from coding & agent
- Muse agent browser blocks pasting, frustrating login flows; 1Password integration requested — altryne · 2026-09-11
- OpenAI Agents API hits public beta; Cloudflare ships sandbox integration for cloud Codex agents — ritakozlov · 2026-09-11
- Greg Isenberg: GPT-6 Astra unlocks physical product startups that needed $2M two years ago — Rasmic · 2026-09-11
- Stress test: running atelier nested inside itself works and stays fast — lucasmeijer · 2026-09-11
- antirez Is Working on DeepSeek V4.1 Support for His ds4 Editor — backyard_tractorbeam · 2026-09-11
- The hard part of agents was never the model — it's hours-long reliability, says dev on OpenAI's Agents API — shaunralston · 2026-09-11