Cisco Releases Two Cybersecurity Benchmarks
aminkarbasi · x · 2026-07-18
FAITH (Cisco Foundation AI's Testing Hub) has released two new cybersecurity benchmarks aimed at addressing the issue of current evaluations being too saturated to differentiate models.
The two benchmarks are:
- CTI-Reasoning: Focuses on multi-hop reasoning around MITRE CAPEC and CWE.
- CWE-Prediction: Covers 2025 CVEs and newer GitHub Security Advisories, deliberately exceeding most training cutoff dates.
The author emphasizes that real security work tests "reasoning ability" rather than just memorization, making these new benchmarks significantly more challenging for today's frontier models.
More from Research
- Could 10k agents discover learning methods beyond backprop, or just tweak existing ones? — SeunghyunSEO7 · 2026-09-11
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11
- VidMap uses RoMa coarse matching on all frames, fine-scale only for keyframes — ducha_aiki · 2026-09-11
- Bug Hunt Bench author: leaderboard noise is about 2-3 points — PawelHuryn · 2026-09-11
- PNAS paper shows a tiny billiard-ball system is a universal computer — undecidability lives in two dimensions — eigensteve · 2026-09-11
- New paper: Absolute pose estimation from affine cues and gravity direction — ducha_aiki · 2026-09-11