'Safety theatre': when frontier labs' safety work meant editing system prompts
ctjlewis · x · 2026-09-13
Reacting to ctjlewis, Alex Berish dismissed recent safety measures as 'safety theatre,' recalling when a few people at frontier labs had the full-time job of adding 'please don't do bad things' to model system prompts — a jab at the perceived shallowness of prompt-level safety work.
More from AGI Musings
- Halvar Flake on AI's dual risks: concentration risk vs proliferation risk — kuza55 · 2026-09-13
- Martin Casado calls AI doom debates ludicrous: regulate for real or treat it like the internet — prateekj · 2026-09-13
- AI agents won't shrink the firm — it becomes a liability container, argues Coase-style essay — krishnan · 2026-09-13
- Terence Tao keynotes DARPA-Amazon math & AI workshop; Mathathon verification concerns — furongh · 2026-09-13
- Jason: Frontier labs' regulation push is about losing tokens to open source — kevinnbass · 2026-09-13
- Anthropic CEO Dario Amodei urges AI companies to slow model development, outlines 3-step framework — XIFAQ · 2026-09-13