Sakana AI’s Fugu-Cyber update tops real-world security benchmarks
SakanaAILabs · x · 2026-07-21
Sakana AI says its new **Fugu-Cyber** update for the Fugu orchestration model reaches state-of-the-art results on real-world security benchmarks. - The company says it matches cyber-focused frontier models such as **GPT-5.5-Cyber** and **Mythos Preview**. - In the attached charts, Fugu-Cyber scores **86.9** on **CyberGym** versus **85.6** for GPT-5.5-Cyber and **83.1** for Mythos Preview. - On **CTI-REALM**, it scores **72.1**, ahead of **GPT-5.5** at **67.3** and **Mythos Preview** at **68.5**.
Related event: Sakana AI Launches Fugu-Cyber Model Achieving SOTA in Security(3 posts)→
More from Models
- Kimi K3 and Fable 5 show nearly identical failure patterns on a software benchmark — FinanceYF5 · 2026-07-21
- Kimi K3 hits 89.4% peak on software tasks while Fable 5 is slightly steadier — FinanceYF5 · 2026-07-21
- Kimi K3 leads on Go, but Fable 5 wins Python, JavaScript, TypeScript and Rust — FinanceYF5 · 2026-07-21
- Kimi K3 costs $4.65 per run and delivers 2.8× more work per dollar than Fable 5 — FinanceYF5 · 2026-07-21
- Kimi K3 reaches 89.4% pass@4 and tops the benchmark over GPT-5.6 Sol — FinanceYF5 · 2026-07-21
- Kimi K3 and Fable 5 now look much closer than the old open-vs-closed gap — FinanceYF5 · 2026-07-21