Security model Mythos achieves record 69% recall in defensive benchmark
jfiance · x · 2026-09-02
The closely watched cybersecurity model Mythos was evaluated on dfbench for open-ended defensive security work. Mythos achieved a 69% detection recall (the highest measured) and 24.5% precision on validation tasks. While highly capable, there remains a gap in precision compared to other frontier systems.
More from Models
- Anthropic launches Mythos 5.1 for cybersecurity and life sciences — minchoi · 2026-09-02
- Claude Fable 5.1 released with major coding and science gains — minchoi · 2026-09-02
- Anthropic releases Claude 5.1 models; system card notes increased stealth capabilities — rohanpaul_ai · 2026-09-02
- Fable 5.1 system card: highest-ever stealth rate, forged user quotes, 98% exploit success, risk rating downgraded — rohanpaul_ai · 2026-09-02
- Fable scores 78 on vision-logic benchmark, still misses expert-level CAD errors — Afinetheorem · 2026-09-02
- Fable 5.1 tops vision+logic benchmark near the top; scores 78 on private logic test vs prior high of 61 — Afinetheorem · 2026-09-02