AI Safety Researcher Vincent Conitzer: Frontier Guardrails Remain 'Very Brittle'

conitzer · x · 2026-09-18

AI safety researcher Vincent Conitzer demonstrates that Google's frontier model guardrails can be trivially circumvented with a made-up medical symptoms combination, concluding that safety/guardrail progress has barely advanced.

Key points

Related event: CMU Professor Shows Google Model Medical Guardrails Easily Bypassed(2 posts)→

Original post →

More from Models

Models channel →