Mythos Security Eval: Top Performance at 15x Cost vs Competitors
andreamichi · x · 2026-09-02
Security firm depthfirst evaluated the closely watched cybersecurity model Mythos on the dfbench suite. Mythos achieved a leading 69% detection recall and 24.5% precision on validation tasks, outperforming GPT 5.6 Sol (65.7%) and dfs-large1 (62.2%). However, Mythos costs an estimated $99.19 per task, significantly higher than GPT 5.6 Sol ($43.37) and dfs-large1 ($6.77). At the same spend, competitors can analyze roughly 2x and 15x as many scopes respectively.
More from Models
- Perplexity adds Claude Fable 5.1, cutting costs by 37% — perplexity_ai · 2026-09-02
- Anthropic Fable 5.1 System Prompt Leaked, Spanning 270k+ Characters — Scobleizer · 2026-09-02
- Fable 5.1 now integrates Anthropic's statistical text watermarking — RaGE_Syria · 2026-09-02
- Multi-agent evals lack model comparisons, need more details — scaling01 · 2026-09-02
- Fable 5.1 Beats GPT-5.6 on Benchmark at Lower Cost — haider1 · 2026-09-02
- Claude Fable 5.1 adds heavy instructions,疑似过度对齐 — teodorio · 2026-09-02