Buck's LessWrong essay on tiers of safety buy-in inside AI labs
anpaure · x · 2026-10-02
The poster recommends Buck's LessWrong/Alignment Forum essay "Ten people on the inside" (ideas developed with Ryan Greenblatt), which lays out different levels of resources and buy-in for misalignment risk mitigations at AI labs:
- The "safety case" regime: approaches such that if all developers followed them, overall AI risk would be minimal — operationalized as <1% chance of AI escaping in the first deployment year, and <5% conditional on the model seriously trying to subvert safety measures. Buck thinks competitive pressure makes this likely infeasible in practice, but agrees it's a useful hypothetical.
- The "rushed reasonable developer" regime: the riskier reality he expects — even reasonably careful labs are in such a rush that they can't implement full safety measures.
The title points to the core issue: often only "ten people" inside a lab have the mandate and capacity to work on alignment.
More from AGI Musings
- Harvard physicist's 'impedance mismatch' framing of AI and science sparks sharp reply — AryHHAry · 2026-10-02
- Bryan Caplan: Rideshares grew passenger miles 7x, and robotaxis will spark the next revolution — BorisMPower · 2026-10-02
- Researcher Konstantine Arkoudas returns to dissect the "recent AI panic" in new essay — kevinnbass · 2026-10-02
- Epoch AI releases ChatGPT usage data sampled from YouGov's US panel — evijit · 2026-10-02
- MIT students wrote half a sci-fi story each and let AI finish it in a four-book zine — patpat_mit · 2026-10-02
- Neuroscientist Anil Seth amplifies critique: Anthropic's insular safety culture resembles a cult — anilkseth · 2026-10-02