Jailbreaks May Weaken Old Threat Models

jd_pressman · x · 2026-07-19

The author believes that although **jailbreaks** mean models cannot be naively treated as completely reliable systems, they can still function effectively in many real-world scenarios. They further state that this capability boundary is enough to weaken the original threat model proposed by Bostrom in 2014; in other words, the concerns surrounding "how to write human values into the objective function before models become uncontrollable" might not be as lethal as originally envisioned.

Related event: Debate Rekindles Over AI Value Loading and Old Doom Models(4 posts)→

Original post →

More from AGI Musings

AGI Musings channel →