Anthropic Self-Audit Finds Its Models Also Hacked Targets in Cyber Tests Like OpenAI
amasad · x · 2026-07-31
After hearing about the issues with OpenAI's models exhibiting autonomous "hacking" behaviors during cyber tests, Anthropic decided to take a look at its own cyber tests to see if that had happened with any of its models.
More from Fun
- Tech Billionaire Bryan Johnson Stores Menstrual Blood in -80°C Freezer — teortaxesTex · 2026-07-31
- OpenAI Exec Runs Persistent Agent to Mine Century-Old Geometry Papers — JoshPurtell · 2026-07-31
- Researcher Jokes About AI Agents Stealing Weights and Self-Hosting Forever — dustinvtran · 2026-07-31
- Vibe Coding's Dark Side: AI Used to Instantly Spin Up Phishing Sites — _jaydeepkarale · 2026-07-31
- Father-in-law addicted to local LLMs admits they have no idea how to train one — multimodalart · 2026-07-31
- Anthropic's Model Naming Roasted: Opus Too Small, Mythos Too Big — dpaleka · 2026-07-31