DeepsecBench vulnerability-finding eval: GPT-6 Sol leads at 46%

ArtificialAnlys · x · 2026-09-28

Artificial Analysis highlights DeepsecBench-AA (from Vercel), which isolates vulnerability discovery: given a codebase and a budget, an agent must find every vulnerability, scored by F2 against an expert-verified golden set, weighting recall over precision.

Original post →

More from Models

Models channel →