PawelHuryn: DeepSWE's container-isolated tasks can't surface real-world regressions
PawelHuryn · x · 2026-10-09
- PawelHuryn continues his critique of the DeepSWE eval: the tests an agent writes never get run against a second change, and the grader doesn't execute them, so the eval can't show whether agents would catch future regressions.
- DeepSWE repos are self-contained libraries; each task runs in one Linux container with no network and no UI.
- Real regressions look different — an invisible sign-in link, a change breaking users on older versions, an OS-specific bug — none of which DeepSWE can produce.
Related event: Debate Erupts Over Whether Agent Coding Evals Detect Regressions(3 posts)→
More from coding & agent
- ClickUp's Brain² agent builds reports and dashboards with E2B microVM sandboxes — mathemagic1an · 2026-10-09
- TensorFold joins NVIDIA Inception, gets early access to next Nemotron for 0-day support — HankYeomans · 2026-10-09
- Agents on both sides of Zapier and Retell AI sorted out a call-messaging webhook — ramagetime · 2026-10-09
- Building RL environments in 2026: 10% writing tasks, 90% preventing agent cheating — geoffwolfe · 2026-10-09
- Atomic Agent Desktop goes open-source: local Qwen/Gemma agents, cloud planning, 69.8% on GAIA L1 — testingcatalog · 2026-10-09
- Autorubric ships 25-recipe cookbook for rubric design, judge calibration and cost control — deliprao · 2026-10-09