DeepMind's Andrew Trask: adversarially train models to break sandboxes until no holes remain

iamtrask · x · 2026-10-08

Google DeepMind senior scientist and OpenMined founder Andrew Trask argues sandboxes are ideal for adversarial training: with an air-gapped cluster, models can repeatedly break and fix a sandbox until no holes remain, preventing cheating (citing the Hugging Face incident). In the full interview he also predicts AI will end as millions of ensembled, per-prompt-routed models beating any single frontier model on quality and price — more like the PC and internet than the mainframe — and discusses why embedded evaluators aren't enough and who really gets to audit AI labs.

Related event: DeepMind Researcher: AI's Future Is Millions of Small Models, Not One Giant(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →