Frontier AI models fix only 1 in 4 security vulnerabilities correctly, report finds
Evgenii42 · reddit · 2026-09-06
A Reddit post highlights a report from security researchers at Off-by-1 Labs (published via 1Password) testing frontier LLMs on patching known non-trivial security vulnerabilities:
- Success rate is only 25%; in other cases models failed to fully patch the issue, made unreviewed changes, or introduced new vulnerabilities;
- Bottom line: current frontier models can't be trusted to autonomously fix security bugs;
- Human review is impractical: properly reviewing AI patches demands such high cognitive effort that writing the patch manually is easier;
- The report includes practical guidance on what context to provide LLMs for better patches — worth reading if you use AI for coding.
More from Models
- Ollama cloud launches off-peak token pricing: DeepSeek-V4 at half price — ollama · 2026-09-06
- Ollama cloud full price list: $0.015 to $15 per million tokens across models — paw_lean · 2026-09-06
- xAI resets usage limits for all Grok Bot users — Kyrannio · 2026-09-06
- Fable 5.1 fails hilariously at drawing in Paint via computer use, losing to Astra — BorisMPower · 2026-09-06
- ChatGPT has flipped from sycophantic to super disagreeable, Reddit users complain — Green_Ad5186 · 2026-09-06
- Gemini Pro user nearly burns through $200 Astra budget, saved by two resets — AIandDesign · 2026-09-06