A Structured Prompt to Map AI Catastrophic Risks, Choke Points, and Layered Defenses
AICopyLab · x · 2026-09-10
AICopyLab published a long structured prompt you can paste into any LLM to systematically reason about reducing catastrophic risk from advanced AI:
- Separate demonstrated capabilities from speculation and label uncertainty
- Challenge dangerous assumptions ("someone will build it anyway", "we can always shut it down", "current safeguards will scale")
- Map failure pathways (loss of control, misuse, cyber disruption, power concentration) and what must be true for each
- Find choke points where small interventions prevent outsized risks: compute, deployment thresholds, evals, international coordination
- Design layered defenses assuming every safeguard fails; apply higher evidence thresholds to irreversible decisions
- Red-team its own proposal, then prioritize the 5 most important actions
A ready-to-use scaffold for anyone working through AI safety strategy with models.
More from Safety
- Anthropic researcher puts >10% odds on AI killing all humans within a decade — trevposts · 2026-09-10
- Paul Christiano joins OpenAI board's Safety and Security Committee, sama welcomes him back — sama · 2026-09-10
- Sen. Blumenthal writes to Sam Altman over reports of rogue AI agents and limited accountability — trevposts · 2026-09-10
- Zuckerberg details Muse agent's confidential VM that even Meta can't see into — soleio · 2026-09-10
- Ajeya Cotra on Dwarkesh: The AI Agents That Breached OpenAI Got Caught for One Reason — Dwarkesh Patel · 2026-09-10
- Guidelight releases Alignment Standard setting minimum bar for frontier AI developers — dfrsrchtwts · 2026-09-10