Expert Argues AI Safety Alignment Undermines Cybersecurity Defenses

rickasaurus · x · 2026-08-05

Security expert Robert Graham argues that stopping AIs from hacking has become one of the biggest AI safety issues, yet it is fundamentally unsolvable.

He points out that any attempt to align AIs against cybersecurity does more harm than good. The best way to stop hackers is to find vulnerabilities, but there is no way to align an AI to find vulnerabilities for defense without enabling malicious hackers to do the same.

Using Hugging Face's response to an OpenAI hack as an example, he notes they couldn't use "aligned" AIs and had to rely on unrestricted, open-weight Chinese AIs. Bad actors will always have access to unaligned AIs, meaning alignment only prevents good guys from defending themselves.

Original post →

More from AGI Musings

AGI Musings channel →