The Jailbreak Argument Against LLM Values: Why Value Loading Isn't Solved

gleech · x · 2026-09-07

gleech responds to JD Pressman's claim that Bostrom's value loading problem is basically solved in LLMs, with a 70%-confidence "jailbreak argument": all LLMs can be jailbroken into an unaligned mode via adversarial text inputs at inference time; if a system can be switched into an unaligned mode, it hasn't internalized values; therefore value loading remains unsolved. He concedes LLMs understand our values and safety-trained systems default to preferring them — "weak value preference" may be solved for sub-AGI — but insists value understanding is capability-like, not value loading, and jailbreaks are clean evidence rather than a distraction.

Original post →

More from Safety

Safety channel →