Cribl Releases SecIT Bench: Diagnostic Investigation Costs Vary 20x Across Models

jonathan_wilke · x · 2026-08-20

Cribl introduced SecIT Bench, a benchmark for evaluating AI agents in real-world IT and security workflows. Testing 14 frontier models across 30 realistic incident scenarios revealed that while diagnostic accuracy had a narrow spread, investigation costs varied roughly 20x. The benchmark aims to provide a rigorous standard for evaluating AI reasoning, performance in real investigations, and cost efficiency.

Related event: Cribl Launches SecIT Bench, Revealing 20x Cost Gap Across AI Models(2 posts)→

Original post →

More from Models

Models channel →