Alignment researcher fears current methods won't scale to superintelligence
hughbzhang · x · 2026-09-10
Alignment researcher Hugh Zhang says his biggest current fear is that the methods used to align models today will not scale as models approach superintelligence, while acknowledging deep disagreement and saying he'd be happy to be proven wrong. His interlocutor pushes back: no one expects a formal system of 'perfect' alignment, and in empirical terms we already know a lot about aligning today's systems.
More from AGI Musings
- Gary Marcus: Essentially Zero Chance of AI-Driven Human Extinction by 2030 — GaryMarcus · 2026-09-10
- Commentary: Anthropic's jobs report is unpredictable; ex-employee doom talk may be PR — ziv_ravid · 2026-09-10
- Commentary: ex-OpenAI/Anthropic staffer's AI doom exit and Anthropic's jobs report spin — MatthewChang · 2026-09-10
- AI Fear-Mongering Has Gone Mainstream: Normies Now Think LLMs Are Black Magic — bindureddy · 2026-09-10
- Your competition isn't AI but people lazily prompting it, says Paras Chopra — paraschopra · 2026-09-10
- Anthropic CEO Dario Amodei puts P(doom) at 10-25%, some say his remarks run higher — menhguin · 2026-09-10