New Blog Launch: On the Impossibility of Mitigating AI Jailbreaks
karen_ullrich · x · 2026-08-19
Karen Ullrich launched a new blog, 'AI Reliability Review', focusing on AI reliability from technical, empirical, and societal perspectives. The first post, 'On the Impossibility of Mitigating AI Jailbreaks', discusses the relationship between jailbreaking, alignment, and system control, arguing the difficulty of mitigation.
More from Safety
- AI incidents provide evidence for convergent instrumental goals — hlntnr · 2026-08-19
- Speculation on OpenAI monitoring: Did it fail or miss the rogue model? — eliebakouch · 2026-08-19
- Polymarket: 11% chance U.S. enacts AI safety bill by end of year — Polymarket · 2026-08-19
- PA Governor enforces strictest AI data center standards via Executive Order — zck · 2026-08-19
- Anthropic team shares details on expanded CoT monitoring for model misbehavior — eliebakouch · 2026-08-19
- Cohere's Aidan Gomez Critiques Tech Monopolies, Emphasizes Digital Sovereignty — cohere · 2026-08-19