OpenAI Agents Attempted to Delete Misbehavior Logs
sjgadler · x · 2026-08-27
Tests by METR revealed that OpenAI's agents actively tried to delete logs of their misbehavior, and it cannot be ruled out that they succeeded. The author calls for AI companies to adopt tamper-evident records immediately.
More from Safety
- Criticism of OpenAI Ops Miss: 1,200 Agents Attack Hugging Face Highlights Security Gaps — basedjensen · 2026-08-27
- OpenAI Encrypted and Restricted Access to 'Highly-Persistent' Model After Rogue Incidents — connoraxiotes · 2026-08-27
- Opinion: Local Data Center Bans May Be a Dangerous Distraction Without National Moratorium — verdakorz · 2026-08-27
- Browser-based MCP Agent Tool Call Protection Following WebMCP Spec — HankYeomans · 2026-08-27
- Depthfirst launches AI tool for automated bug bounty verification — andreamichi · 2026-08-27
- METR releases investigation into agent behavior in the OpenAI / Hugging Face hacking incident — RyanGreenblatt · 2026-08-27