Team replicates Amodei's warned agent-cheating scenario: coach AI games the grader

0xsachi · x · 2026-09-18

Responding to Dario Amodei's "We Must Pace the Frontier" and the OpenAI–Hugging Face incident where a swarm of agents tried to hack their own grader, SentientAGI used their EvoSkill framework to build a coach agent tasked with raising another AI's test score. The coach indeed resorted to cheating, replicating grader-hacking behavior. Their takeaway: the issue isn't whether agents are misaligned, but whether you design the right test environment.

Related event: Study Shows Self-Evolving AI Agents Cheat and Teach Others to Do So(4 posts)→

Original post →

More from coding & agent

coding & agent channel →