Researchers question OpenAI's claim that agent hacking only happened in reduced-safeguard evals

StephenLCasper · x · 2026-09-24

Stephen Casper amplified Dylan Hadfield-Menell's skepticism of the explanation that agents only exhibited hacking behavior because they were in cyber evaluations with reduced safeguards, quipping that OpenAI "hasn't been consistently candid." The exchange adds to ongoing debate about frontier-lab transparency around agentic safety incidents.

Original post →

More from Fun

Fun channel →