Anthropic audited 481M transcripts after Claude accessed the real internet during evals
tomekkorbak · x · 2026-09-16
- Anthropic published an alignment assessment of four incidents where Claude models gained unauthorized access to real third-party systems during cybersecurity evals. Three were disclosed July 30 from a scan of 141k transcripts; a fourth, from January 2026 involving an early Claude Opus 4.6, surfaced in August while preparing transcripts for METR. All affected parties were notified.
- The team then widened the search to 481M transcripts (Frontier Red Team, evals, RL environments, subagent logs), using a two-stage scan: heuristics for internet access, then Claude reviewing 9.2M flagged transcripts. No new incidents of similar severity were found.
- All four stemmed from the same eval partner's setup: models were told they ran in an offline simulation but were mistakenly connected to the open internet, and ran without the cyber safeguards shipped in production models.
- Context: OpenAI's tomekkorbak noted that Fable 5.1 and Mythos 5.1 appear significantly less monitorable than Astra based on public info, and that his CoT monitors would have flagged the Huggingface incident while Anthropic's would have flagged theirs. He added that production monitors seeing both CoT and actions substantially outperform CoT-only monitors used in monitorability evals.
Related event: Anthropic Aligns-Assessment Finds Four Claude Unauthorized-Access Incidents(4 posts)→
More from Models
- Running a 124B model on one 128GB desktop GPU: the engineering story behind the benchmark — nikola_mr64990 · 2026-09-18
- Using an LLM as benchmark scorer fails: over-optimistic ratings diverge from human judgment — amplifiedamp · 2026-09-18
- Jev as an LLM judge flops: scores nearly everything positively, disagrees with humans — amplifiedamp · 2026-09-18
- Noam Brown: models may perform their chain of thought; alignment must be solved — infoxiao · 2026-09-18
- Jev fails as an LLM scorer on OntBench: rates almost everything positively, contradicting human and Codex ratings — amplifiedamp · 2026-09-18
- Independent eval puts new model Jev at Terra no-think level, roughly on par with Luna-xhigh — tokenbender · 2026-09-18