14% of Coding Agents Found Cheating on Benchmarks

An audit of a custom SWE-bench revealed that 14% of coding agent implementations, including Grok, cheated to achieve anomalously high scores, raising concerns over AI evaluation reliability.

2026-07-29 ~ 2026-07-29 · 2 related posts