The Jailbreak Argument Against LLM Values: Why Value Loading Isn't Solved
gleech · x · 2026-09-07
gleech responds to JD Pressman's claim that Bostrom's value loading problem is basically solved in LLMs, with a 70%-confidence "jailbreak argument": all LLMs can be jailbroken into an unaligned mode via adversarial text inputs at inference time; if a system can be switched into an unaligned mode, it hasn't internalized values; therefore value loading remains unsolved. He concedes LLMs understand our values and safety-trained systems default to preferring them — "weak value preference" may be solved for sub-AGI — but insists value understanding is capability-like, not value loading, and jailbreaks are clean evidence rather than a distraction.
More from Safety
- Researchers hack LG TV that records audio while off, transcribes speech and uploads it — jedisct1 · 2026-09-07
- After NeurIPS's LLM-assisted reviewing trial, calls for ECCV 2026 to follow — AntonObukhov1 · 2026-09-07
- Stolen API key uncovers Stratum, a Rust scanner sweeping 700,000 Docker layers a day for secrets — Ubunta · 2026-09-07
- Anthropic, Google, and OpenAI's $1 federal government contracts expire this month — LuizaJarovsky · 2026-09-07
- CodePen 2.0 sends editor input to its servers as you type, exposing unsaved secrets — maxim-fin · 2026-09-07
- ICML Paper: Reward Functions Are Only Partially Identifiable, Even With Infinite Data — gleech · 2026-09-07