Mythos Security Eval: Top Performance at 15x Cost vs Competitors

andreamichi · x · 2026-09-02

Security firm depthfirst evaluated the closely watched cybersecurity model Mythos on the dfbench suite. Mythos achieved a leading 69% detection recall and 24.5% precision on validation tasks, outperforming GPT 5.6 Sol (65.7%) and dfs-large1 (62.2%). However, Mythos costs an estimated $99.19 per task, significantly higher than GPT 5.6 Sol ($43.37) and dfs-large1 ($6.77). At the same spend, competitors can analyze roughly 2x and 15x as many scopes respectively.

Related event: Security Model Mythos Hits Record 69% Recall but at High Cost, Low Precision(2 posts)→

Original post →

More from Models

Models channel →