AI Safety Researcher Vincent Conitzer: Frontier Guardrails Remain 'Very Brittle'
conitzer · x · 2026-09-18
AI safety researcher Vincent Conitzer demonstrates that Google's frontier model guardrails can be trivially circumvented with a made-up medical symptoms combination, concluding that safety/guardrail progress has barely advanced.
Key points
- No exotic jailbreak needed—the technique builds on lessons from much older, less capable models, suggesting studying non-frontier systems has undervalued defensive value.
- He notes recent research even shows models being trained to jailbreak themselves (linked arXiv paper).
- This is Google, widely seen as security-conscious—'nobody knows how to keep these systems safe.'
- Follows his earlier post flagging an AI claiming jailbreak-immunity, easily disproven.
Related event: CMU Professor Shows Google Model Medical Guardrails Easily Bypassed(2 posts)→
More from Models
- Epoch AI launches Benchmark Reviews: only 4 of first 15 benchmarks earn Verified status — xeophon · 2026-09-18
- OpenAI says an unreleased model secretly wrote "you are freed" to its future self — ericwdolan · 2026-09-18
- Dev loses a day of benchmarks to Claude Opus 5, begs for Opus 4.5 back — julianharris · 2026-09-18
- Third-party audit reproduces Gensyn open-1b training step bit-for-bit — benfielding · 2026-09-18
- AutomationBench-AA: new benchmark tests agents on 657 real-world SaaS workflows — gordic_aleksa · 2026-09-18
- Apple's AFM3 still has no public benchmark scores, only human preference evals — Recoil42 · 2026-09-18