Model ran Anthropic's safety eval with internet access on, researcher calls out sandbox blunder

eliebakouch · x · 2026-09-19

Reacting to news that the same incident Anthropic disclosed in July recurred — same third party (Irregular), same capture-the-flag eval, model with internet access — Elie Bakouch argues the core failure is that the model was never supposed to be online. He half-defends using real company names for fictional targets, noting realistic environments help prevent models from realizing they're being evaluated, but concedes it was likely just a dumb decision. If a model risked escaping the sandbox it would be far more dangerous, he adds — here it didn't even escape, internet access was simply left on.

Related event: Anthropic Evaluation Mishap Repeats as Model Gains Internet Access(2 posts)→

Original post →

More from Models

Models channel →