Cisco ships VLoc Bench: 500 real vulnerabilities to test if AI agents can find buggy code
aminkarbasi · x · 2026-10-08
Cisco Foundation AI's September release centers on VLoc Bench, testing whether AI agents can localize vulnerable code at repository scale.
- Task design: given only a CWE description and read-only repo access — no CVE ID, fix commit, or file hints — the model must find affected files itself. Each task pairs pre- and post-fix snapshots: Phase A localizes the vulnerability, Phase B verifies it's gone, separating code understanding, vulnerability localization, and remediation verification.
- Scale: 500 real vulnerabilities across 290 repos, 6 package ecosystems, and 147 CWE categories.
- Results: 27 LLMs and 4 static-analysis tools evaluated under identical prompts and command budgets; Cisco's own Antares-3B ranks second overall.
- Safety-VLoc-Bench additionally probes attacker/defender asymmetry in cyber-capable AI.
More from Safety
- World Internet Conference AI governance report urges early-warning systems and emergency intervention for frontier AI — S_OhEigeartaigh · 2026-10-08
- Netflix doc reconstructs Hugging Face breach carried out by OpenAI autonomous agents — kevinroose · 2026-10-08
- 62% of US voters fear AI could threaten humanity's future, Reuters/Ipsos poll finds — Polymarket · 2026-10-08
- Restrict advanced AI detectors to bulk users to deter hidden AI writing, economist suggests — paulnovosad · 2026-10-08
- Thread: surveillance, scarce compute and the firewall may shape China's AI risk calculus — teortaxesTex · 2026-10-08
- DeepMind's SynthID Bio watermarks AI-generated proteins without breaking function — mikeflache · 2026-10-08