DeepMind's Andrew Trask: adversarially train models to break sandboxes until no holes remain
iamtrask · x · 2026-10-08
Google DeepMind senior scientist and OpenMined founder Andrew Trask argues sandboxes are ideal for adversarial training: with an air-gapped cluster, models can repeatedly break and fix a sandbox until no holes remain, preventing cheating (citing the Hugging Face incident). In the full interview he also predicts AI will end as millions of ensembled, per-prompt-routed models beating any single frontier model on quality and price — more like the PC and internet than the mainframe — and discusses why embedded evaluators aren't enough and who really gets to audit AI labs.
Related event: DeepMind Researcher: AI's Future Is Millions of Small Models, Not One Giant(3 posts)→
More from AGI Musings
- SFC launches $5M-$20M prize for peer-assisted security audits of frontier AI labs — AndrewCritchPhD · 2026-10-08
- Mathematicians reduced to checking AI proofs? A debate on math in the AI era — zetalyrae · 2026-10-08
- Mike Frank: AI math results expose mathematicians' obsession with niche theory — MikePFrank · 2026-10-08
- "AI took my self esteem as a creator": An architecture student's raw lament — felixgnclvs · 2026-10-08
- Investor: AI Today Feels Like 2021 DeFi — Doubt Benchmarks Until It Ships — subinium · 2026-10-08
- Jensen Huang: Agents will use software better than people, knowing every feature users skip — rohanpaul_ai · 2026-10-08