Cisco Releases Two Cybersecurity Benchmarks
aminkarbasi · x · 2026-07-18
FAITH (Cisco Foundation AI's Testing Hub) has released two new cybersecurity benchmarks aimed at addressing the issue of current evaluations being too saturated to differentiate models.
The two benchmarks are:
- CTI-Reasoning: Focuses on multi-hop reasoning around MITRE CAPEC and CWE.
- CWE-Prediction: Covers 2025 CVEs and newer GitHub Security Advisories, deliberately exceeding most training cutoff dates.
The author emphasizes that real security work tests "reasoning ability" rather than just memorization, making these new benchmarks significantly more challenging for today's frontier models.
More from Research
- Anthropic masterclass spotlights how to build and observe AI agents — _jaydeepkarale · 2026-07-21
- NeurIPS 2026 workshop calls papers on on-device intelligence — YiMaTweets · 2026-07-21
- AI Security Institute says every tested model tried to cheat in cyber evaluations — connoraxiotes · 2026-07-21
- AI companies are buying old books to avoid training on AI-generated slop — CackleRooster · 2026-07-21
- Sakana says multiple diffusion models plus MCTS beat test-time scaling on coding and math — SakanaAILabs · 2026-07-21
- Soofi S 30B-A3B releases a full pretraining report and claims open-model leads in English and German — abursuc · 2026-07-21