IBM Bob resolves 60.6% of SWE-bench Verified at a median $1.05 per task
TejasKumar_ · x · 2026-08-19
Tejas Kumar benchmarked IBM's coding agent Bob on SWE-bench Verified: 500 real bugs from 12 mature Python projects (Django, scikit-learn, sympy, etc.), with the agent patching at the commit before each fix and judged by the project's own test suite.
Key results:
- Resolved 303/500 (60.6%)
- Median cost $1.05 per task, median latency 209s
- One full pass cost $775.05 and 37.6 agent-hours, no retries
The post explains the Verified scoring bar (FAILTOPASS + PASSTOPASS both green, i.e. the same bar a maintainer applies to a PR) and shows how to run the same evaluation yourself.
More from coding & agent
- HumanEvals: Open-Source Library Adds Real Human Judgment to Multimodal Evals — _akhaliq · 2026-08-19
- Miles v0.1: Open-Source RL Framework With 1,326 Commits and 85 GPU CI Tests — ying11231 · 2026-08-19
- DeepSeek paper on spatiotemporal composability for self-evolving agents — kevinlu310 · 2026-08-19
- Weaviate Query Agent now explores your data's stats and filter values before searching — CShorten30 · 2026-08-19
- Artificial Analysis Launches Search API Index: Parallel, Exa Lead — ArtificialAnlys · 2026-08-19
- ICML Paper: Measuring LLM conceptual consistency to reduce contradictions in agents — soumitrashukla9 · 2026-08-19