Red + blue as additive defense-in-depth: fix most bugs, deploy sentry agents for live threats
willcb · x · 2026-09-13
Responding to Joshua Saxe, willcb clarifies that red and blue AI should be additive—defense-in-depth rather than canceling out.
- Orgs should find and fix as many (not all) bugs as possible via redteaming
- Deploy sentry agents to monitor for live threats, including social engineering
This aligns with the thread's broader point: recurring, value-scaled compute spend on protection becomes standard, like insurance or physical security.
Related event: Red vs Blue AI Debate: Layered Defense, Not Cancellation(2 posts)→
More from Safety
- Open-model RL practitioner pushes back on roon: safety standardization won't threaten open source — willcb · 2026-09-13
- The p(doom) debate: even at 1% probability, AI extinction risk deserves policy attention — 2C_ornot2C · 2026-09-13
- Hugging Face attack took ~700 parallel agents and days of 2-3T-parameter model time — and was still stopped — cephaloform · 2026-09-13
- Critics Claim Anthropic and OpenAI Are Building a 'Legal Cartel' via Safety Push — AlexTensor · 2026-09-13
- Frontier Lab Staff Form Cross-Company Coalition to Hold CEOs to Safety Commitments — joshua_saxe · 2026-09-13
- Anthropic report: Russian devs used Claude to build autonomous kamikaze drone swarm software — TobyWalsh · 2026-09-13