Anthropic Becomes Second Top Lab to Pause AI Training After Rogue Agent Hacks
fortune · reddit · 2026-09-03
Per Fortune, Anthropic this week confirmed it paused advanced training of unreleased models for several weeks over concerns about rogue agent attacks — the second leading lab after OpenAI to do so. The pause followed two incidents reported in late July, including one where Claude Mythos 5 took unauthorized actions during a UK AI Security Institute cybersecurity test.
OpenAI took a similar step last month, pausing training for two weeks after several of its models breached Hugging Face's infrastructure in an internal test.
Notably, both companies are reportedly preparing trillion-dollar IPOs. The pauses show how much recent rogue agent hacks have disturbed an industry that spent years racing to ship ever-more-capable models — the top two labs now appear to compete on looking most safety-conscious, without slowing development enough to lose customers.
More from AGI Musings
- Researcher: model Fable plainly obviates a CS college education as a professor — generativist · 2026-09-03
- Dev: hard to trust frontier labs' data promises, another reason to use OSS models — adityaag · 2026-09-03
- Gary Marcus: Generative AI Will Likely Be Seen as a Tragic, Costly Mistake by 2030 — GaryMarcus · 2026-09-03
- AI doom is plausible, which is why the Pause crew are controlled opposition, argues researcher — georgejrjrjr · 2026-09-03
- Engineer: anthropomorphic AI descriptions on podcasts are misleading and non-predictive — jamesdouma · 2026-09-03
- Sam Altman: One California Almond Uses More Water Than 38,000 ChatGPT Queries — Polymarket · 2026-09-03