AI safety theatre critique: LLMs are unsafe by design and need external controls

gerardsans · x · 2026-09-14

A pointed take arguing the AI industry has long engaged in safety theatre: prompts have no safety boundaries, since instructions, CoT, external data, observations and tools all share one context, with facts, lies and jailbreaks one tool call away — unsafe by design.

The author's core argument: you can't fix this with more instructions or gradient descent, because an LLM is not a mind but a mathematical object, a soft program. The path forward is deterministic, auditable external controls rather than more alignment discourse about a mind that doesn't exist.

Original post →

More from Safety

Safety channel →