Grok 4.7 posts 59% recall on defensive cyber bench at half the cost of rivals
andreamichi · x · 2026-09-22
Depth First Labs got early access to Grok 4.7's red-team capabilities for defensive security research and benchmarked it on dfbench.
- Solid mix of recall and precision on defensive cyber tasks: 59% recall and 23.9% precision
- Cost advantage is significant: roughly half the cost of GPT 5.6 Sol and one-fifth the cost of Mythos 5
- Reviewer calls it a good cyber model with fast progress in this area
More from Models
- Xiaomi's MiMo-V2.6-Pro debuts as top open-weights model with 46 on AA Intelligence Index — huggingface · 2026-09-22
- Open multilingual System 1 decision model tops Hugging Face trending — huggingface · 2026-09-22
- Independent Tests Show Grok 4.6 (high) Beating 4.7, 23 vs. 19 — PawelHuryn · 2026-09-22
- mimo-v2.6-pro Claimed to Redraw the Price-Performance Pareto Frontier — zainhas · 2026-09-22
- Unverified: Grok 4.7 out now, costs more per task than GPT, says leaker — ChrisGPT · 2026-09-22
- MiMo v2.6-Flash-RL vs open-weight peers: community chart fills the missing comparison — ababaka · 2026-09-22