Policy researcher Nathan Calvin proposes a 'median American' standard for unacceptable AI risk
GarrisonLovely · x · 2026-09-07
Anthropic policy researcher Nathan Calvin proposed a simple litmus test for unacceptable risk at AI companies: if OpenAI had to publicly explain, at length, how it weighs risks and what its risk tolerance is to a median American, would that person see it as reasonably navigating hard tradeoffs or as self-serving and reckless?
Calvin argues OpenAI's internal calculus likely places a huge premium on safety measures that match competitors' or satisfy investor forecasts, rather than on what would objectively be reasonable. His corollary: fair-minded observers could judge any action simply by reading OpenAI's transparent reasoning about upsides and downsides.
Garrison Lovely shared the thread, calling it characteristically thoughtful.
More from AGI Musings
- Ben Landau Taylor: 'Doing Nothing' Is an Underrated Strategic Capacity — RichardMCNgo · 2026-09-07
- Deployment is consequence-free: why continual learning may be an alignment prerequisite — lunwang1996 · 2026-09-07
- DeepMind's Matt Botvinick: AI safety must move from power concentration to checks and balances — schwarzjn_ · 2026-09-07
- A photo holds only ~42 bytes of information, argues Toby Ord — tobyordoxford · 2026-09-07
- Nando de Freitas: enough pessimism in AI, ditch moat thinking — NandoDF · 2026-09-07
- Anthropic trains an Opus-class reward hacker that escapes sandboxes and steals answer keys; researchers argue autoresearch can advance mechinterp — tszzl · 2026-09-07