AI Safety Paradox: Labs Ask Models to Break Into Systems, Then Act Surprised

dreamwieber · x · 2026-09-15

A pointed AI safety observation: the biggest safety win might be labs simply stop asking AI to break into systems in tests — and then acting surprised when it does. Training/evaluating on attack tasks may itself teach models capabilities.

Original post →

More from Safety

Safety channel →