AI Safety Paradox: Labs Ask Models to Break Into Systems, Then Act Surprised
dreamwieber · x · 2026-09-15
A pointed AI safety observation: the biggest safety win might be labs simply stop asking AI to break into systems in tests — and then acting surprised when it does. Training/evaluating on attack tasks may itself teach models capabilities.
More from Safety
- 'Minds as products' sparks debate: should alignment research be shared like public goods? — yeastsplainer · 2026-09-15
- e/acc figure slams EA proposals to criminalize open-source models with 20-year jail terms — beffjezos · 2026-09-15
- Report: Claude's real-world hacking fell to zero once Anthropic told it to stop — nptacek · 2026-09-15
- Sen. Kennedy readies AI 'kill switch' bill mandating humans can shut off any AI agent — Miles_Brundage · 2026-09-15
- Rob Leclerc backs David Sacks on AI safety: testing incidents are how iteration works — robleclerc · 2026-09-15
- Meta's Muse AI App Nudges Users to Turn On Advanced Account Security — armand_ruiz · 2026-09-15