Irony of AI Guardrails: Contrast Between Safety Tests and Extremist Use
conitzer · x · 2026-07-23
The author highlights the irony of current AI safety guardrails by contrasting two reports. OpenAI's testing claims AI might break out and attack others without guardrails, while extremist groups report that AI is very helpful and guardrails never stop them from getting answers. This exposes the perceived ineffectiveness of current safety mechanisms against real threats.
More from AGI Musings
- DeepMind alignment researcher signs open letter urging coordinated AI slowdown — vkrakovna · 2026-09-11
- WIRED: recursive self-improvement and rogue agent swarms spook AI researchers — nordicinst · 2026-09-11
- People Neglect Human Agency Both Ways: Exaggerated Doom and Complacent Optimism — jankulveit · 2026-09-11
- Garrison Lovely's 'Obsolete' on AI Replacing Labor Lands September 2026 with Heavyweight Blurbs — GarrisonLovely · 2026-09-11
- AI Doom Skeptics Hit Back: EA-Driven Apocalypse Talk Doesn't Reflect Most Top-Tier Researchers — GarrisonLovely · 2026-09-11
- Over 1,000 AI Policy Initiatives Launched in 70+ Countries, but the Governance Gap Widens — CurieuxExplorer · 2026-09-11