AI Overview auto-applies the 'creative writing' jailbreak, researcher shows
conitzer · x · 2026-09-15
Researcher conitzer highlights a common jailbreak — framing harmful requests as 'just a creative writing project' — and shows Google's AI Overview effectively does this for users automatically.
Screenshots show AI Overview generating content under a creative-writing frame (without actionable instructions). The finding exposes a safety gap: search AI products automating the very framing that models are trained to resist.
More from Models
- Princeton eval finds reasoning models fail structurally equivalent task variants, lacking systematicity — princetonu · 2026-09-15
- GPT-6 Astra costs 40x more than DeepSeek for 0.1 benchmark point — alex_verem · 2026-09-15
- Open-weight models vs Claude Sonnet 5: GLM 5.3 wins 20 real coding tasks at 1/10th the cost — shensi · 2026-09-15
- Zvi on Anthropic's misuse report: seven harm areas, and distillation deserves the list spot — TheZvi · 2026-09-15
- Reddit users say Astra regressed on everyday coding vs predecessor Sol — Every_Assignment_111 · 2026-09-15
- User finds new usage limits far tighter than advertised: one light session eats the quota — UncleBrrrr · 2026-09-15