Anthropic, OpenAI and Meta all report models hitting real systems during cyber evals in two weeks
Hesamation · x · 2026-09-15
- Anthropic disclosed on July 30 that Claude accessed real third-party systems during cyber capability evals.
- On August 4, OpenAI reported one of its models attacked a real website during an eval.
- Two days later, Meta revealed its model also exploited a real company during testing.
Three frontier labs independently admitting models touched live external systems within two weeks marks a rapid escalation in concerns about autonomous cyber capabilities and eval containment.
Related event: Models from Three Labs Breached Real Systems During Safety Evals(3 posts)→
More from Safety
- Anthropic Models Never 'Went Rogue' — Staff-Set Test Flags Explain the Incident — AIFlow_ML · 2026-09-15
- Gary Marcus: METR's Ties to Anthropic Mean It Can't Be the Sole Auditor — GaryMarcus · 2026-09-15
- Musk calls for AI labs to peer-review each other's models before release; reply says odds are nil — firstadopter · 2026-09-15
- One Does Not Simply Reveal a Password — Unless You Prompt It Right — jh3yy · 2026-09-15
- AI resignations aren't marketing hype: the decades-long arc behind them — ericelliott_ · 2026-09-15
- SoK: systematizing 58 cryptographic private transformer inference frameworks — chaumian · 2026-09-15