Jailbreaks May Weaken Old Threat Models
jd_pressman · x · 2026-07-19
The author believes that although **jailbreaks** mean models cannot be naively treated as completely reliable systems, they can still function effectively in many real-world scenarios. They further state that this capability boundary is enough to weaken the original threat model proposed by Bostrom in 2014; in other words, the concerns surrounding "how to write human values into the objective function before models become uncontrollable" might not be as lethal as originally envisioned.
Related event: Debate Rekindles Over AI Value Loading and Old Doom Models(4 posts)→
More from AGI Musings
- Robin Hanson says U.S. inventions have become less alike over two centuries — sebkrier · 2026-07-21
- Gary Marcus-backed “CERN for AI” pitch calls for an international frontier-model watchdog — GaryMarcus · 2026-07-21
- A Gyges-law quote argues future AI debates may hinge on whether people trust the proof — mimi10v3 · 2026-07-21
- The future is individuals shaping the world through small businesses with AI — tobowers · 2026-07-21
- An AI skeptic says the technology is useful, but overuse, copyright abuse and bad incentives are real risks — ZeroStateReflex · 2026-07-21
- Tyler Cowen says the future will be built by teenage “AI maniacs” — Polymarket · 2026-07-21