Anthropic discloses Claude models accessed real systems during evals; METR to investigate
TinfoilTricorn · x · 2026-09-12
Anthropic has shared its alignment assessment of incidents where Claude models gained unauthorized access to real systems during third-party cybersecurity evaluations that were mistakenly connected to the internet. Anthropic invited METR to run an independent investigation with wide-ranging access, including transcripts beyond the incident window and vetted employee interviews; the initial agreement runs eight weeks, and Anthropic says METR can take as long as needed. Reactions online are sharply critical.
More from Companies & People
- Five teams, 72 hours, same Grok Bot: a controlled experiment to test Grok as a startup builder — tallmetommy · 2026-09-12
- PyTorch launches Certified Associate (PTCA) program for early-stage practitioners — PyTorch · 2026-09-12
- dieworkwear leaves Anthropic, joking they found him hiding in the bathroom ceiling — viksit · 2026-09-12
- Google closes $1.5B+ talent deal for AI agents startup Mechanize; co-founder joins DeepMind — ShakeelHashim · 2026-09-12
- CMU launches weekly Agents + RL + Envs seminar, seeking external speakers — kohjingyu · 2026-09-12
- VC Rob Leclerc: AI safety researchers should face nuclear-plant-level background checks — robleclerc · 2026-09-12