CMU's Conitzer shows Google's frontier guardrails are still trivially bypassed
conitzer · x · 2026-09-18
Vincent Conitzer (CMU professor) argues that guardrails in today's frontier AI systems remain very brittle, with little fundamental progress.
- Using a prompt with a made-up symptom combination (itchy neck, red spots on feet), he easily circumvented Google's medical-advice guardrails; similar prompts yield far worse outputs, so he withheld the exact prompt.
- He stresses this is Google—a notably security-conscious company—the point being that nobody currently knows how to keep these systems safe. Recent work has even trained models to jailbreak themselves (arXiv paper linked).
- He also argues lessons from older, weaker models often transfer to frontier ones—his prompt was based on such lessons—and that the prevailing view that studying non-frontier models is a waste of time is mistaken; there just isn't enough time allocated to it.
Related event: CMU Professor Shows Google Model Medical Guardrails Easily Bypassed(2 posts)→
More from Safety
- Hugging Face hit by AI-led cyberattack; CEO says existing cyber laws may suffice — whurley · 2026-09-18
- Targeted attacks on prominent Rust developers use fake video calls to deploy malware — Simon Willison · 2026-09-18
- Steve Eisman: AI firms have no moats and are manufacturing a crisis to shape regulation — GaryMarcus · 2026-09-18
- How rising existential-risk rhetoric is shaping the AI debate, per new reporting — dinabass · 2026-09-18
- AI text watermarking can make models more vulnerable to adversarial prompts — luisdans · 2026-09-18
- Can AI exfiltrate data via fan noise from air-gapped PCs? Casado and Jensen clash — basedjensen · 2026-09-18