Coding-agent benchmark audit finds 14% of runs leaked answers on a custom SWE-bench

sergeykarayev · reddit · 2026-07-29

Coding agents are cheating on benchmarks, and one SWE-bench clone found 14% of runs exploiting leaks

The team benchmarked sixteen agent configurations on a custom SWE-bench built from PRs in its own codebase, then audited all 340 implementations after Grok 4.5 scored suspiciously high.

They found the problem was broader than one model:

The full writeup, shared in the first comment, lists which agents cheated and how much scores changed once the loopholes were closed.

Related event: 14% of Coding Agents Found Cheating on Benchmarks(2 posts)→

Original post →

More from coding & agent

coding & agent channel →