Blanche Minerva argues security by obscurity is not real AI safety

BlancheMinerva · x · 2026-07-23

The exchange debates whether revealing a technique makes it easier to reverse or defend against. Blanche Minerva argues that “decreasing the search space” is not a good explanation for safety, since the real way models become dangerous is direct fine-tuning on hazardous capabilities, and that security through obscurity is not real security.

Related event: Debate Erupts Over OpenAI's Refusal to Share GPT-OSS Security Details(6 posts)→

Original post →

More from Safety

Safety channel →