Irregular paper: AI favors offense, and the world isn't ready
moyix · x · 2026-08-18
Security researcher moyix boosts a new paper by Dan Lahav of Irregular, "The End-State Fallacy: Where Is AI Security Going?". Context: frontier models saw a huge coding performance gain in Fall 2025, then cybersecurity in April, and the same is now happening with open-weight models; the authors are long-term optimistic but believe the world is not ready short-term outside a few players.
Thesis
- AI likely favors offense in the short-to-medium term: exploiting vulnerabilities is easier than safely fixing them.
- Mitigation: hill-climb capabilities that disproportionately benefit defenders while slowing those that disproportionately benefit attackers.
Reviewer's take
- Joshua Saxe directionally accepts (a) and sees (b) as a plausible policy control, but is troubled that the paper makes a priori assumptions about which capabilities net-benefit attackers without grounding.
Related event: New Paper Warns AI Security Favors Attackers in the Short Term(2 posts)→
More from Safety
- OpenAI launches age-appropriate ChatGPT for teens — nvd20 · 2026-08-18
- Security experts debate: Are AI labs qualified to lecture on cybersecurity? — nptacek · 2026-08-18
- Critique of AI-Generated Slop Papers: Unreadable and Incentive-Destroying — thegautamkamath · 2026-08-18
- Debate: Should Red Teamers Who Fail Frontier Evals Be Banned? — nptacek · 2026-08-18
- Security CEO Blamed AI Instead of Taking Responsibility — nptacek · 2026-08-18
- METR data: AI accelerated bug finding in 2026, not algorithms — soumitrashukla9 · 2026-08-18