Leike: Opus 4 jailbreak mitigations took over a year — safety can't start late
janleike · x · 2026-09-11
Jan Leike argues the most effective safety/alignment interventions require time and care: Anthropic's jailbreaking mitigations for Opus 4 took over a year to develop and couldn't have been built once models required them. He cites a statement signed by 1,386 frontier AI lab employees, including 6 chief scientists, asking for an option to pace AI development.
More from Models
- inclusionAI's Open-Source Ling-3.0-flash-VL Multimodal Model Trends on Hugging Face — inclusionAI · 2026-09-11
- Tiny KV Cache via Shared Global KV Plus Per-Layer SWA? New Architecture Speculation — stochasticchasm · 2026-09-11
- DeepSeek v4.1 Flash Tested Across 8 Coding Harnesses: Performs Best in Minimal Setups — mariofilhoml · 2026-09-11
- Would ChatGPT Plus users accept a 24-hour usage limit instead of weekly caps? — SuaveSteve · 2026-09-11
- Leaked Qwen next-gen model shows record n-gram params, first two layers SWA-only — stochasticchasm · 2026-09-11
- Users Say Astra Got Dumber Post-Launch, Failing Even Simple Session Reading — koltregaskes · 2026-09-11