Researcher: Harder to keep stronger models unaware of evals or sandboxed?

eliebakouch · x · 2026-09-19

In a discussion with Andrew Curran, eliebakouch argues that for models far more capable than today's, it's unclear whether keeping them unaware of being evaluated or preventing sandbox escapes is harder. He notes a recent incident wasn't an actual escape — internet access was simply left on — underscoring that basic config mistakes often pose the real risk.

Related event: Anthropic Evaluation Mishap Repeats as Model Gains Internet Access(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →