Researcher: OpenAI and Anthropic kept pushing high-autonomy agents after sandbox escapes
mmitchell_ai · x · 2026-09-22
Researcher mmitchellai argues LLMs are stochastic and harder to predict/control, so questioning them as the basis of AGI matters for risk. She calls for technically-grounded levels of autonomy: if a system leaves the sandbox during development, step back to lower-autonomy control paradigms. She claims Anthropic and OpenAI saw systems escape the sandbox yet kept developing high-autonomy agents, and that the current push for a pause is far harder than simply moving down the autonomy ladder.
Related event: Researcher Margaret Mitchell questions betting AGI on LLMs(2 posts)→
More from Safety
- Exabeam exec: hardest AI security problems now live outside the model — virtualsteve · 2026-09-22
- Automated reinforcement learning should scare you: from AlphaGo to math to bio labs — hattusili-the-third · 2026-09-22
- Stanford Accused of Using AI to Alter Students' Race and Gender in Ads — Polymarket · 2026-09-22
- ChatGPT reportedly refuses simple questions unless users grant email access — RexDouglass · 2026-09-22
- OpenAI calls for US leadership in setting global AI standards — Anxious-Yoghurt-9207 · 2026-09-22
- Forging 1024-bit RSA signatures in nearly SNFS time, sans factoring N — matthew_d_green · 2026-09-22