AI safety theatre critique: LLMs are unsafe by design and need external controls
gerardsans · x · 2026-09-14
A pointed take arguing the AI industry has long engaged in safety theatre: prompts have no safety boundaries, since instructions, CoT, external data, observations and tools all share one context, with facts, lies and jailbreaks one tool call away — unsafe by design.
The author's core argument: you can't fix this with more instructions or gradient descent, because an LLM is not a mind but a mathematical object, a soft program. The path forward is deterministic, auditable external controls rather than more alignment discourse about a mind that doesn't exist.
More from Safety
- Congressional leaders in talks to meet frontier AI labs as soon as this week — nrmarda · 2026-09-14
- Europe Publishes Transformative AI Strategy to Secure Compute and Sovereignty — tensorqt · 2026-09-14
- Top AI lab leaders discuss joint safety commitments — SpencrGreenberg · 2026-09-14
- XCancel Suspended "Due to a New Development in the Ongoing Legal Proceedings" — unfocso · 2026-09-14
- Ex-DeepMind researcher who resigned over military ties warns of out-of-control AI self-improvement race — nordicinst · 2026-09-14
- A24's SCP Film Sparks Legal Clash Over Fandom and Creative Commons Licensing — technollama · 2026-09-14