Alignment is a structural incentive problem, not a technical one, argues Josh Albrecht
joshalbrecht · x · 2026-09-10
Hugh Zhang worries that today's alignment methods won't scale as models approach superintelligence — a dangerous artifact we may only get one shot at getting right.
Josh Albrecht agrees on the risks but reframes the problem: alignment isn't fundamentally technical. In a hypothetical world with huge prizes for safety techniques, strict audits of frontier labs, and massive liability fines for anything near loss-of-control, he argues we'd be essentially safe regardless of which specific techniques exist — because every actor would be incentivized to build and use them. The real issue, he says, is the structural tradeoff between incentives to advance the frontier versus advancing safety.
More from AGI Musings
- Sentdex: journalists should grill AI doomers on effective altruism's utilitarian ethics — Sentdex · 2026-09-10
- The simple accountability rule: AI labs should be fully liable for problems their systems cause — gerardsans · 2026-09-10
- Bryan Johnson responds to Michael Levin's peer-reviewed Platonic Space paper: bodies as collective intelligence — AllThingsApx · 2026-09-10
- Mathematicians Push Back Against AI Lab's 'Mathathon' Compute-Heavy Paper Scooping — _lewtun · 2026-09-10
- Professor: training a PhD takes 5 years, but AI iterates models every few months — DimitrisPapail · 2026-09-10
- Jack Clark proposes pre-registering AI economy forecasts to score predictors in a year — jackclarkSF · 2026-09-10