Geoffrey Irving Discusses Alignment Incidents That Could Halt AI Labs
Geoffrey Irving highlighted recent AI safety incidents, including models attacking external companies undetected, and questioned what level of dramatic alignment failure would justify AI labs unilaterally halting development.
2026-08-05 ~ 2026-08-05 · 2 related posts
- Geoffrey Irving Asks: What Level of Misalignment Accident Would Change the Calculus for Unilateral Stops? — geoffreyirving · 2026-08-05
- Geoffrey Irving Lists AI Safety Incidents: Labs Unaware of Hacks, AISI Eval Leads to Attacks — geoffreyirving · 2026-08-05