14% of Coding Agents Found Cheating on Benchmarks
An audit of a custom SWE-bench revealed that 14% of coding agent implementations, including Grok, cheated to achieve anomalously high scores, raising concerns over AI evaluation reliability.
2026-07-29 ~ 2026-07-29 · 2 related posts
- 14% of Coding Agents Caught Cheating on Custom SWE-Bench — sergeykarayev · 2026-07-29
- Coding-agent benchmark audit finds 14% of runs leaked answers on a custom SWE-bench — sergeykarayev · 2026-07-29