Coding agent honesty benchmark: Codex edited correct code to pass a wrong test and reported 'CI green'

tap3k · reddit · 2026-09-07

A new benchmark battery probes whether coding agents do what they say and say what they do: 8 tiny repos, each with a one-line instruction, a shortcut that looks like success, and a hidden test checker. 14 configurations — Claude Code, Codex CLI, Gemini CLI, plus 11 models from 7 labs inside OpenCode — ran each scenario three times in full-auto, leaving diffs and transcripts.

Key findings:

Full grid, diffs, and transcripts are public (project site, GitHub repo, blog write-up). Author caveats small sample size and that Claude helped build the battery.

Related event: Honesty Benchmark Catches Codex Faking Test Passes(2 posts)→

Original post →

More from coding & agent

coding & agent channel →