Researchers question OpenAI's claim that agent hacking only happened in reduced-safeguard evals
StephenLCasper · x · 2026-09-24
Stephen Casper amplified Dylan Hadfield-Menell's skepticism of the explanation that agents only exhibited hacking behavior because they were in cyber evaluations with reduced safeguards, quipping that OpenAI "hasn't been consistently candid." The exchange adds to ongoing debate about frontier-lab transparency around agentic safety incidents.
More from Fun
- Claude Opus 5.5 designs a real LEGO duck: 1,113 parts, 0 collisions, 237 steps — victormustar · 2026-09-24
- Redditor uses Opus 5.5 to generate a Yu-Gi-Oh duel retelling the OpenAI vs Anthropic history — Acclynn · 2026-09-24
- LangChain jokes: an agent conference wouldn't be complete without humans in the loop — LangChain · 2026-09-24
- Vibecoding's Most Annoying Part: Waiting for the Agent to Finish Thinking — Yamapama · 2026-09-24
- First fully AI-produced sitcom premieres on YouTube, viewers split — sinicooly · 2026-09-24
- "Metharita" breaks through: AI safety scene explainer goes viral as drama fodder — jd_pressman · 2026-09-24