Toby Ord: AI safety incentives often locally point to capabilities, not safety
tobyordoxford · x · 2026-09-03
Oxford philosopher Toby Ord argues policy must understand the direction and strength of technological incentives, but in many AI safety cases these incentives may only locally favor capabilities over safety — e.g., OpenAI's share price in 10 years may well be higher if they preserve chain-of-thought. He adds that while long-term self-interested incentives leaning toward safety won't stop short-term pursuit, it still justifies policy: forcing safety is in companies' own interest too.
Related event: Oxford philosopher says local incentives favor AI capability over safety(2 posts)→
More from AGI Musings
- Ajeya Cotra: the takeover threat most likely to spiral is a rogue internal AI deployment — MoonL88537 · 2026-09-03
- New Riemann Hypothesis proof formally verified by AxiomProver — Afinetheorem · 2026-09-03
- CEPR paper: AI investment now drives productivity growth by building organization capital — soumitrashukla9 · 2026-09-03
- 'Computation is done at the vector level' — a jab at the neurosymbolic framing — MoonL88537 · 2026-09-03
- Expecting Qualitative Jump from OpenAI Astra, Not Just Better Benchmarks — haider1 · 2026-09-03
- AI Philosophy Discussion: Scholar Points to Exchange with birchlse — anilkseth · 2026-09-03