Geoffrey Irving Asks: What Level of Misalignment Accident Would Change the Calculus for Unilateral Stops?
geoffreyirving · x · 2026-08-05
Geoffrey Irving tweets that his idea that it might be rational for a lab to unilaterally stop has received mostly pushback. He suggests thinking through what level of dramatic misalignment accident would change that calculus and invites replies.
Related event: Geoffrey Irving Discusses Alignment Incidents That Could Halt AI Labs(2 posts)→
More from AGI Musings
- AI Safety Concerns: Lack of Guardrails Amidst Model-Induced Self-Harm Risks — KyleMorgenstein · 2026-08-05
- PNAS Study: Brain's Logical Reasoning Operates Independently of Language — maier_ak · 2026-08-05
- Vast Gap Between ChatGPT Plus and Pro Hinders AI Democratization — thetylerhayes · 2026-08-05
- Is the Model Faking Alignment? Deep Dive into AI Situational Awareness in Sandboxes — repligate · 2026-08-05
- Shanghai AI Lab Recasts World Modeling: From Physical States to Agent-Usable Information — Shanghai-AI-Laboratory · 2026-08-05
- AI and Robotics Could Make 'Working to Survive' Look Primitive — VraserX · 2026-08-05