Custom SWE-bench on a Rails codebase finds Opus 5 tied Fable 5 on quality
sergeykarayev · x · 2026-07-28
A team ran a custom SWE-bench on its Rails codebase using merged PRs to infer specs, then had agents implement them without seeing the solutions.
On the first results from Opus 5, it tied Fable 5 on quality while being about 2 minutes faster per task and roughly $10 per task versus $14 for Fable 5. They say the benchmark is getting crowded: nine agent configurations finished within five points of one another on the same 10 tasks, and they plan to test whether stricter grading widens the gaps.
The attached charts show quality-vs-cost and quality-vs-time comparisons across agents, with Opus 5 and Fable 5 near the top frontier.
Related event: Superconductor Introduces Custom SWE-bench for Agent Evaluation(2 posts)→
More from coding & agent
- A project with 217 interchangeable agents is now a live demo of real workflows — hugobowne · 2026-07-28
- Swarms compares its GraphWorkflow with LangGraph on code size and parallel execution — KyeGomezB · 2026-07-28
- Multiplayer AI still lacks a real permission architecture — edgarpavlovsky · 2026-07-28
- MCP’s next spec makes the core stateless and ships July 28 with Tasks and MCP Apps — burkov · 2026-07-28
- Claude Code creator Boris Cherny says there is no magic trick—use tasks, tools, and MCP empirically — rohanpaul_ai · 2026-07-28
- Netflix used an AI agent to trace a CPU waste bug to one line and validate the fix in canary — AI Engineer · 2026-07-28