1Password study: AI fixes only 26% of security bugs, unreviewed LLM patches are net-negative

GaryMarcus · x · 2026-09-06

Researchers at 1Password ran over 6,000 AI-generated patches from Claude and ChatGPT against real vulnerabilities, with sobering results: AI fixed the security bug only 26% of the time, failed to fix the original bug half the time, and introduced brand-new vulnerabilities 4.5% of the time.

They also found that reviewing an AI-generated patch often takes more effort than writing the fix yourself, concluding bluntly that "the expected value of a fully LLM-generated, non-human-reviewed patch is a net-negative by a considerable margin."

Gary Marcus shared the study while criticizing OpenAI's "Patch the Planet" program, which aims to have AI find and fix bugs autonomously.

Related event: Study Finds Frontier AI Fixes Only 26% of Security Vulnerabilities(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →