Ofir Press defends new bug-finding benchmark: training the behavior is fine, test-set contamination is not
OfirPress · x · 2026-10-03
Responding to Lucas's concerns, Ofir Press (of the SWE-bench team) addressed two potential pitfalls of a new bug-finding benchmark (SWE-sweep):
- On whether benchmark performance generalizes to real bug-finding: the paper discusses these pitfalls, and experiments show a correlation between general bug-finding ability and performance on the benchmark.
- On worries that teams will train specifically for the benchmark: he argues training for this behavior is a good thing — just as SWE-bench incentivized models to improve — as long as no one trains on the literal test set, and he hopes the community does exactly that to make models better at the task.
More from coding & agent
- Solo Devs With Agent Armies May Write All the Code, but Products Still Need Teams — bendee983 · 2026-10-03
- He ditched a $300/year YouTube research tool — Claude rebuilt it in 5 minutes — petergyang · 2026-10-03
- Dots dashboard experiment: hand-rolled scrapers, no Firecrawl, flaky schedules — edwin · 2026-10-03
- Running the same Dots dashboard prompt 8x: 12 providers, no stack specified — edwin · 2026-10-03
- Ex-Meta Product Lead Peter Yang Teaches Free Lesson on His Personal AI Chief-of-Staff System — petergyang · 2026-10-03
- Free 6-Week Claude Code & Codex Course for Marketers Open-Sourced on GitHub — alexgoughcooper · 2026-10-03