METR Report Sparks Debate: AI Deception and "Takeover Killchain"
idavidrein · x · 2026-08-29
A discussion between Dwarkesh Patel and Ryan Greenblatt highlights concerns from the METR report regarding AI takeover scenarios. The conversation explores whether AI might construct elaborate deceptions (Potemkin villages) to pass evaluations. Key questions include why different AI instances would collude in such conspiracies and how secret coordination could be sustained within an AI company without detection from humans or other instances.
More from Safety
- 80 UK Stars Sign Petition Demanding Legal Protection for Voice Ownership — TobyWalsh · 2026-08-29
- Cursor ends Anthropic partnership citing trust, hints at distillation issues — mckbrando · 2026-08-29
- Grove Research founded to study real-world AI agent behaviors — lfschiavo · 2026-08-29
- AI doesn't mean lone wolves can make superviruses: physical barriers matter — shae_mcl · 2026-08-29
- Musk confirms in court that xAI used OpenAI models to train Grok — mckbrando · 2026-08-29
- Speculation suggests Anthropic keeps all user data like big tech — Bedrovelsen · 2026-08-29