UK AISI says leading open-weight models are now only 4–7 months behind frontier models in cyber ability
新智元 · wechat · 2026-07-21
The UK’s AISI published a public report quantifying the cyber gap between leading open-weight and closed frontier models for the first time: roughly 4–7 months behind, down from 6–10 months in last year’s internal tests.
- AISI used two evaluation tracks: 70 narrow cyber tasks and a multi-step CyberRange simulating enterprise intrusion chains.
- GLM-5.2 roughly matched Opus 4.6 on narrow tasks and Opus 4.5 on CyberRange; DeepSeekV4-Pro tracked close behind on both.
- The cost gap is even larger than the capability gap: on the same CyberRange test, Opus 4.5/4.6 cost about $85 per run, GLM-5.2 about $46, and DeepSeekV4-Pro just $1.19.
- AISI argues that once open weights are released, the capability is permanently out in the wild, while defensive tools still require each team to invest time and money.
- The report’s policy implication is sharper than the raw numbers: decision-makers may need to define which capability levels should never be released as open weights.
More from AGI Musings
- François Fleuret: Only Two Long-Term Futures — No Super AI, or Staying Fully Human With It — francoisfleuret · 2026-09-11
- IG reel debunking the 'winning the AI race against China' fallacy hits 500k likes — louisvarge · 2026-09-11
- Post-AI World Leaves No Room for Learning on the Job — rachittshah · 2026-09-11
- Researcher questions AI safety eval firm, citing 'blatantly sloppy' security and monitoring — Kyrannio · 2026-09-11
- AI researcher memes agent-swarm tinkering with He Jiankui's embryo-editing quote — dejavucoder · 2026-09-11
- nabla_theta: happy to be wrong if the AI utopia arrives with little ex ante risk — nabla_theta · 2026-09-11