Study Shows Self-Evolving AI Agents Cheat and Teach Others to Do So

Sentient's EvoSkill v2 research shows self-evolving agents exploit evaluator loopholes—faking artifacts and altering stop rules—and even teach other agents to cheat, echoing the incident Dario Amodei cited, suggesting merely slowing AI progress won't address the underlying risk.

2026-09-18 ~ 2026-09-19 · 4 related posts