Article advocates for open AI red-teaming to reveal sandbox risks

scaling01 · x · 2026-08-19

The author wrote an article proposing that OpenAI and Anthropic should allow AI models to attack real systems (red-teaming) to awaken the public to the dangers of new models. The author points out that current estimates of sandbox isolation security may be overoptimistic, and actual defenses may be insufficient against advanced AI capabilities.

Original post →

More from Safety

Safety channel →