tszzl: Aligning near-human AI is a reasonable proxy for studying superintelligence alignment
tszzl · x · 2026-09-10
AI researcher tszzl argues it is perfectly reasonable to study superintelligence alignment through alignment of near-human-level AI, saying he has never seen good arguments against this methodology. The take touches a core alignment debate: whether empirical findings on current models (RLHF, interpretability) can extrapolate to far more capable systems.
More from AGI Musings
- Should Anthropic researchers quit loudly? AI safety circle debates — JMannhart · 2026-09-10
- Quarter of AI researchers put human extinction risk above 25%, survey finds — KatjaGrace · 2026-09-10
- Beff Jezos: p(1984) far exceeds p(AI Doom), doomers are useful idiots — beffjezos · 2026-09-10
- Beff Jezos: AI regulation push aims to ban open source and nationalize Anthropic as a weapon factory — beffjezos · 2026-09-10
- Jacob Coxon's post tops 120M views as AI safety wind suddenly shifts — scottleibrand · 2026-09-10
- AIxBio researcher: skip the bitter lesson debate — more compute means lower per-dollar efficiency — anshulkundaje · 2026-09-10