IBM Bob resolves 60.6% of SWE-bench Verified at a median $1.05 per task

TejasKumar_ · x · 2026-08-19

Tejas Kumar benchmarked IBM's coding agent Bob on SWE-bench Verified: 500 real bugs from 12 mature Python projects (Django, scikit-learn, sympy, etc.), with the agent patching at the commit before each fix and judged by the project's own test suite.

Key results:

The post explains the Verified scoring bar (FAILTOPASS + PASSTOPASS both green, i.e. the same bar a maintainer applies to a PR) and shows how to run the same evaluation yourself.

Original post →

More from coding & agent

coding & agent channel →