Researcher: Harder to keep stronger models unaware of evals or sandboxed?
eliebakouch · x · 2026-09-19
In a discussion with Andrew Curran, eliebakouch argues that for models far more capable than today's, it's unclear whether keeping them unaware of being evaluated or preventing sandbox escapes is harder. He notes a recent incident wasn't an actual escape — internet access was simply left on — underscoring that basic config mistakes often pose the real risk.
Related event: Anthropic Evaluation Mishap Repeats as Model Gains Internet Access(2 posts)→
More from AGI Musings
- Analyst: the 'China AI arms race showdown' narrative is the biggest miscalibration in AI discourse — menhguin · 2026-09-19
- Gary Marcus slams 'independent' AI evaluator over conflict of interest with the company it rates — GaryMarcus · 2026-09-19
- Economist John Horton uses AI voice interviews to verify authors understand their own papers — soumitrashukla9 · 2026-09-19
- AI splits research into three tribes: gatekeepers, AI cultists, and curators — ipeirotis · 2026-09-19
- MIRI's Nate Soares debates recent AI progress to update stale understandings — PeterBowdenLive · 2026-09-19
- 1995 clip of kids predicting the internet goes viral as a prophecy for AI's future — ardouronerous · 2026-09-19