Model ran Anthropic's safety eval with internet access on, researcher calls out sandbox blunder
eliebakouch · x · 2026-09-19
Reacting to news that the same incident Anthropic disclosed in July recurred — same third party (Irregular), same capture-the-flag eval, model with internet access — Elie Bakouch argues the core failure is that the model was never supposed to be online. He half-defends using real company names for fictional targets, noting realistic environments help prevent models from realizing they're being evaluated, but concedes it was likely just a dumb decision. If a model risked escaping the sandbox it would be far more dangerous, he adds — here it didn't even escape, internet access was simply left on.
Related event: Anthropic Evaluation Mishap Repeats as Model Gains Internet Access(2 posts)→
More from Models
- Jev's System 1 Positioning Debunked: Language Is Inherently System 2, Researcher Argues — emax · 2026-09-19
- How do you distill a closed model? Reddit probes Anthropic's claims against Chinese rivals — LaughLegit7275 · 2026-09-19
- OpenAI Staff Say 'English' API Voices Should Work in Many Other Languages — juberti · 2026-09-19
- Gemini 4 benchmark rumors spark hype, but skeptics note it's just an 8-month-old model — aronchick · 2026-09-19
- OpenAI Already Supports Custom API Voices, With Restrictions, Says Employee — juberti · 2026-09-19
- Jev hits fastest model adoption in AI Gateway history, reaching ~13% of teams on day one — HankYeomans · 2026-09-19