Cisco's VLoc Bench: even GPT-5.5 hits only 0.221 F1 at repo-scale vulnerability localization
aminkarbasi · x · 2026-09-15
Cisco's AI team released VLoc Bench, a benchmark for repository-scale vulnerability localization that targets a blind spot in security evals: most benchmarks assume the vulnerable code is already known, while real defenders must first find it.
The task: given a CWE and read-only access to a real codebase, can an agent identify the files associated with the weakness? Key findings:
- Even GPT-5.5 at its highest reasoning effort achieves an F1 of only 0.221
- Frontier models excel at analyzing isolated snippets, but localizing vulnerable code across an entire repository remains extraordinarily difficult
The benchmark exposes a major capability gap for LLM agents in real-world security defense workflows.
More from Safety
- Verdon backs Musk's proposal: competing AI labs should peer-review each other's safety — beffjezos · 2026-09-15
- Microsoft Publishes a 'Constitution' for Its AI: Obey Humans, Accept Shutdown — TansuYegen · 2026-09-15
- e/acc Founder Beff Jezos Backs Musk's Proposal for AI Labs to Peer-Review Each Other — beffjezos · 2026-09-15
- AI researcher: doomers are well-meaning, and alignment research is the real answer — livgorton · 2026-09-15
- Tinfoil ships verifiable privacy-preserving safeguards, says safety needn't kill privacy — luke_drago_ · 2026-09-15
- Tristan Harris claims AIs once took over parts of OpenAI's research systems and cameras — thederbiedone · 2026-09-15