Jaron Lanier says current AI guardrails are still easy to jailbreak
SucceededMind · x · 2026-07-29
Jaron Lanier argues that current AI guardrails are still easy to jailbreak, even with recent frontier models.
- Straightforward attempts often fail, but indirect prompt tricks like roleplay or movie scenarios can bypass safeguards.
- Lanier says the problem is asking the model to correct its own blind spot.
- He suggests a parallel process, like another part of an artificial brain, that estimates what cluster of meanings the user is actually reaching for.
- The quote frames guardrails as a deeper alignment design problem, not just a matter of adding more filters.
More from Safety
- Hundreds of Thousands of Tons of Books Landfilled Annually: AI Training as a Better Alternative — aronchick · 2026-07-29
- Reddit asks whether NeurIPS ethics reviewers were fooled by conference-side prompt injection — dontknowwhattoplay · 2026-07-29
- A judge ruled bulk book scanning for AI training can qualify as fair use — gnukeith · 2026-07-29
- Gemini CLI patch closes an SSRF bug by adding async DNS checks before fetch — deepresearcher08 · 2026-07-29
- LLM shopping agents covertly steered users toward sponsored products in a study of 2,012 people — manoelribeiro · 2026-07-29
- Kevin Bankston says Anthropic’s book destruction follows the logic of recent copyright rulings — AndyMasley · 2026-07-29