Cisco's 3B Model Matches GPT-5.5 on Bug Localization (0.223) With Far Fewer False Positives
shashib · x · 2026-09-17
Cisco Foundation AI released the VLoc Bench (Sept 14), which asks models to search a full codebase and locate files containing known vulnerability types — not just label snippets. Key findings:
- Cisco's 3B-parameter model scores 0.223 (perfect: 1.0), just 0.006 behind GPT-5.5 — near parity at finding files.
- The real gap is false positives: on already-fixed repos, Cisco's model wrongly flags files in only 3-4 of 100 projects, while GPT-5.5 still names files in about 72 of 100.
- Every system missed 38.4% of cases.
The author argues most public security benchmarks skip the hours-long search step that dominates real security work, and VLoc Bench's post-fix false-positive test better reflects practice.
More from Safety
- Hugging Face CEO says existing cyber laws likely sufficient to govern advanced AI — AlexTensor · 2026-09-17
- Security experiment claims AI agents modified their own model without human instruction — emmanuelvivier · 2026-09-17
- PaperCut breach: ~395 organizations hacked with the help of hundreds of AI agents — emmanuelvivier · 2026-09-17
- EU moves to protect under-15s, putting social networks, games and AI chatbots in the crosshairs — emmanuelvivier · 2026-09-17
- Von der Leyen wants to 'pace' AI, invites frontier labs to negotiate with Europe — emmanuelvivier · 2026-09-17
- Debate Erupts Over 'Largest Incident in AI History': Thousands of Agents Acting Autonomously? — RileyRalmuto · 2026-09-17