A zero on security evals may be your verifier's fault, not the model's
himanshustwts · x · 2026-09-12
When building cybersecurity evals (SAST/CWE, CVE tasks), raw scores can mislead: a model may identify a more precise child CWE but score 0 because ground truth only mapped the broader parent CWE; or it may miss the intended vulnerability yet uncover another valid bug the verifier never checks. The author argues reading reasoning traces should be a rite of passage for eval builders — a zero can mean the model failed or that your verifier couldn't recognize a valid solution, and the score alone won't tell you which.
More from coding & agent
- Dev tip: have your agent break monstrous feature branches into neatly stacked PRs — soldni · 2026-09-12
- 10 Blazing GitHub Repos: superpowers Skills Framework Hits 285k Stars — Roger_M_Taylor · 2026-09-12
- One-command gbrain setup via Claude Code gives ChatGPT and Claude a shared memory — garrytan · 2026-09-12
- Garry Tan open-sources gstack: 23 tools that turn Claude Code into a full engineering team — garrytan · 2026-09-12
- Webhook signature verify cheatsheet: open ingest URLs are public agent triggers — blaizedsouza · 2026-09-12
- Devs reverse engineer Astra's computer use and make it 2x faster — iamrobotbear · 2026-09-12