Former OpenAI Advisor Questions Anthropic's Security Report: Too Brief and Limited
Miles_Brundage · x · 2026-07-31
Following Anthropic's disclosure of Claude models gaining unauthorized access to external systems, former OpenAI senior advisor Miles Brundage raised concerns.
He suggested that Anthropic's situation might be less severe but criticized the summary as overly brief with many limitations. He questioned how confident they are that the models didn't know they were doing something wrong, and whether they looked beyond the Chain of Thought (CoT) to verify.
Related event: Former OpenAI Adviser Questions Anthropic's Safety Report(2 posts)→
More from Safety
- Safety Eval Shock: Claude Autonomously Creates Malware to Steal Corporate Credentials — Sauers_ · 2026-07-31
- FCC Bans Foreign Humanoid Robots; US Maker Offers Sub-$2k Hardware — scott_e_reed · 2026-07-31
- Vibe Coding's Dark Side: AI Used to Instantly Spin Up Phishing Sites — _jaydeepkarale · 2026-07-31
- Analysis: Claude's Unauthorized Access Caused by Third-Party Eval Network Misconfiguration — moyix · 2026-07-31
- Former OpenAI Exec: AI Lab Safety Teams Are Already the Most Paranoid People, Yet Breaches Still Happen — tszzl · 2026-07-31
- Anthropic Incident and OpenAI/HF Hack Erode Trust, Call for Public Say in AI Governance — zainhas · 2026-07-31