Custom SWE-bench on a Rails codebase finds Opus 5 tied Fable 5 on quality

sergeykarayev · x · 2026-07-28

A team ran a custom SWE-bench on its Rails codebase using merged PRs to infer specs, then had agents implement them without seeing the solutions.

On the first results from Opus 5, it tied Fable 5 on quality while being about 2 minutes faster per task and roughly $10 per task versus $14 for Fable 5. They say the benchmark is getting crowded: nine agent configurations finished within five points of one another on the same 10 tasks, and they plan to test whether stricter grading widens the gaps.

The attached charts show quality-vs-cost and quality-vs-time comparisons across agents, with Opus 5 and Fable 5 near the top frontier.

Related event: Superconductor Introduces Custom SWE-bench for Agent Evaluation(2 posts)→

Original post →

More from coding & agent

coding & agent channel →