CheckerBench: Best Coding Agent Scores Just 45.33% on Static-Analysis Checker Synthesis
humanlaya-data-lab · hf · 2026-10-08
humanlaya-data-lab on Hugging Face released CheckerBench, an executable benchmark targeting long-horizon agent work on static-analysis checker synthesis — a task existing coding-agent benchmarks don't cover.
- Tasks: 300 tasks derived from 297 CVEs across 167 repositories, 85 CWEs, and five language ecosystems; each includes vulnerable/fixed revisions, a pinned analysis environment, and a checker scaffold.
- Evaluation: The companion CheckerLab framework independently rebuilds submitted checkers and measures vulnerable-fixed diagnostic contrast, patch localization, false positives, and tool use.
- Results: Across 21 model-harness configurations with three repeats each, mean Pass@1 is 32.30%, with the best at 45.33%.
The takeaway: building reliable, reusable checkers remains a hard open problem for current coding agents.
More from coding & agent
- Jeffrey Emanuel's "say no to process" agent skill kills Codex ceremony output — used hundreds of times a day — doodlestein · 2026-10-08
- a16z backs Preference Model, which open-sources Karotte RL environment framework battle-tested by 1M+ evals — a16z · 2026-10-08
- Every's agent skims meeting notes and only pings you when your name comes up — here's the 4-step setup — every · 2026-10-08
- Exa's setup page swaps dev docs for a copy-paste prompt your coding agent runs — josh_bickett · 2026-10-08
- Haiku 5.5 targets high-volume tasks, works as a coding subagent with Opus/Sonnet — claudeai · 2026-10-08
- Non-coder runs his entire business on an army of Claude Opus 5.5 agents — EXM7777 · 2026-10-08