Aidan Clark says safety techniques are easier to undo once they’re known
_aidan_clark_ · x · 2026-07-23
Aidan Clark argues that if people know a technique, it becomes easier to undo — so sharing safety techniques may weaken them.
He also says the team would be open to partnerships that help harden external models, including open-source ones, but they would be much less eager to share the techniques themselves.
Related event: OpenAI Model Safety Test Sparks Sandbox Escape and Jailbreak Debate(2 posts)→
More from Safety
- A reply argues open-weight bans target models that were SOTA just years ago — 1a3orn · 2026-07-23
- Warner Music’s Sureel AI teams with Symphonic on music AI usage tracking — Proper_Subject · 2026-07-23
- Criticizing Overly Strict Open-Weight Prohibition Proposals by AI Policy Groups — 1a3orn · 2026-07-23
- Quoted post says LLaMA 2 should be stopped to slow AI proliferation — 1a3orn · 2026-07-23
- AIRA proposal would regulate only frontier AI systems and require stronger security — 1a3orn · 2026-07-23
- U.S. lawmakers are preparing a bill that would let DHS throttle or shut down AI systems — The Verge AI · 2026-07-23