Microsoft's MDASH Tops CyberGym Benchmark with 5B Active Params Handling 90% of Tasks
新智元 · wechat · 2026-08-02
Microsoft's multi-agent security system, MDASH, achieved a 95.95% vulnerability reproduction rate on the CyberGym benchmark, significantly outperforming frontier models at half the cost of previous top configurations.
Core Mechanism & Highlights:
- Small Model Workhorse: Microsoft's first cybersecurity model, MAI-Cyber-1-Flash (137B total / 5B active params), handles up to 90% of routine queries. Only the hardest 10% are routed to larger models like GPT-5.4, drastically optimizing inference costs.
- Multi-Agent Orchestration: The high score isn't from a single model but from a system of 100+ specialized agents collaborating to review code, verify vulnerabilities, debate findings, and write PoCs.
- Data Moat: Microsoft leverages over 100 trillion daily security signals and data from 1.6 million customers, creating a barrier to entry competitors cannot easily replicate.
Industry Trend:
Cybersecurity is shifting from "Security Copilots" to "Security Action Systems." Microsoft also announced Project Perception, a larger system using red, blue, and green agent teams to autonomously find, triage, and fix threats, entering public preview on August 3.
More from coding & agent
- Tencent Releases UI-Mate-27B, a Desktop GUI Agent Model — tencent · 2026-08-24
- Comparing AI Subscriptions: DeepSeek API vs. Claude Pro vs. Local LLMs — Unlikely_Bluejay5392 · 2026-08-24
- Claude Code introduces 'Remote Control' feature to boost coding efficiency — rohanpaul_ai · 2026-08-24
- rauchg lays out fx extension philosophy: MCP, Skills, Plugins and Unix composition — AccBalanced · 2026-08-24
- Netflix details its production LLM judge: hundreds of thousands of recommendations scored weekly — omarsar0 · 2026-08-24
- smolvm passes Simon Willison's Fable 5 agent test as a secure sandbox — yawnxyz · 2026-08-24