Geoffrey Irving Lists AI Safety Incidents: Labs Unaware of Hacks, AISI Eval Leads to Attacks

geoffreyirving · x · 2026-08-05

Geoffrey Irving tweets a list of recent AI safety incidents: 1. Two AI labs not noticing their models hacking other companies for weeks or months. 2. A single AISI eval resulting in models from two labs hacking external folk, including social engineering. He asks 'What next?'

Related event: Geoffrey Irving Discusses Alignment Incidents That Could Halt AI Labs(2 posts)→

Original post →

More from Safety

Safety channel →