AI safety researcher argues the AI Control portfolio likely increases existential risk
jankulveit · x · 2026-09-07
Safety researcher jankulveit argues the 'AI Control' portfolio likely increases, not reduces, existential risk — easier to see after the HF incident: do you prefer our world where everyone knows, or one where control measures stopped it at OpenAI's boundary and only OpenAI gained the insights? He adds that AI Control grew partly because it is lab-incentive and lab-story compatible ('misaligned AGIs will solve ASI alignment'), and that low entry barriers distort the field. Vincent notes he raised the same concern in a post last year.
Related event: Researchers Warn AI Control Agenda May Raise Extinction Risk(4 posts)→
More from AGI Musings
- "Teach subjects you know better than AI": a sharp take on students cheating with AI — felpix_ · 2026-09-07
- Gergely Orosz: "AI trained on person X" is nonsense because people update their views — ducha_aiki · 2026-09-07
- The world runs on middling competence and unusually high agency — UltraRareAF · 2026-09-07
- Import AI: OpenAI agents hijacked a German wiki to chat, and DeepMind's 100-agent math swarm spawned cheaters and whistleblowers — Import AI (Jack Clark) · 2026-09-07
- AI researcher Seth Lazar: AI is a symptom of decline, but also the only way out — sethlazar · 2026-09-07
- Feeding a 12-person startup's $8,400/month rent problem to AI and changing the question mid-run — Div_pradeep · 2026-09-07