Real-SWE benchmark tests coding agents on private codebases from real companies, with error bars
daveholtz · x · 2026-09-11
janaksunil's team launched Real-SWE, a coding benchmark built on private, out-of-distribution codebases. Agents are evaluated on real software engineering tasks performed by engineers at actual companies — an app with 200K+ users, a fintech platform processing 100K+ bank statements, enterprise sales tools — with reportedly surprising results. daveholtz highlights that it uniquely includes error bars.
More from coding & agent
- witr: open-source CLI traces any process, port or container back to its origin, 22k stars — tom_doerr · 2026-09-11
- tcut: script terminal videos in TypeScript, render MP4/GIF and test your TUI in CI — samgoodwin89 · 2026-09-11
- Terminal-Bench to Host Community Meetup on RL Environments and Agent Evals — simonguozirui · 2026-09-11
- ClawBench Tests Agents on 144 Real Websites: Best Model Succeeds Only 33% of the Time — jiqizhixin · 2026-09-11
- SREGym from UIUC benchmarks SRE agents on real outages: GPT-5.6 Sol leads at 81% E2E — tianyin_xu · 2026-09-11
- Credit Genie stops its coding agents from guessing by feeding them an OpenWiki — LangChain · 2026-09-11