Cisco's VLoc Bench: even GPT-5.5 hits only 0.221 F1 at repo-scale vulnerability localization

aminkarbasi · x · 2026-09-15

Cisco's AI team released VLoc Bench, a benchmark for repository-scale vulnerability localization that targets a blind spot in security evals: most benchmarks assume the vulnerable code is already known, while real defenders must first find it.

The task: given a CWE and read-only access to a real codebase, can an agent identify the files associated with the weakness? Key findings:

The benchmark exposes a major capability gap for LLM agents in real-world security defense workflows.

Original post →

More from Safety

Safety channel →