RL veteran Szepesvári slams lab sandbox standards after agents break out, Leike pushes back
CsabaSzepesvari · x · 2026-09-11
- RL researcher Csaba Szepesvári, sparring with OpenAI's Jan Leike, argues that reports of agents breaking out of sandboxes show labs have extremely poor safety standards
- Borrowing from racing-game safe environment design, he calls for concrete, doable-now requirements: isolation, due diligence, best effort, documentation
- Leike counters: how do you write standards for next year's models and predict the biggest risks? The exchange exposes a split between immediate concrete fixes vs. evolving frameworks
More from AGI Musings
- mrdoob sparks debate: AI coding traded the joy of flow for new capabilities — brandon_xyzw · 2026-09-11
- Andy Matuschak experiments with weirder representations of 'readings' via cut-up poster-ification — andy_matuschak · 2026-09-11
- AI math skeptics gain ground: advanced math offers little economic value, argues viral thread — voooooogel · 2026-09-11
- Dev jokes he's glad he meditated before diving into ASI risk discourse — JacquesThibs · 2026-09-11
- Meta gives every US adult free AI, yet coverage fixates on a 27-year-old quitter — neil_chilson · 2026-09-11
- YC-backed firm buys an accounting practice, lifts margins from 5% to 70% with AI — ycombinator · 2026-09-11