UK AISI says leading open-weight models are now only 4–7 months behind frontier models in cyber ability

新智元 · wechat · 2026-07-21

The UK’s AISI published a public report quantifying the cyber gap between leading open-weight and closed frontier models for the first time: roughly **4–7 months** behind, down from **6–10 months** in last year’s internal tests. - AISI used two evaluation tracks: 70 narrow cyber tasks and a multi-step **CyberRange** simulating enterprise intrusion chains. - **GLM-5.2** roughly matched **Opus 4.6** on narrow tasks and **Opus 4.5** on CyberRange; **DeepSeekV4-Pro** tracked close behind on both. - The cost gap is even larger than the capability gap: on the same CyberRange test, **Opus 4.5/4.6** cost about **$85** per run, **GLM-5.2** about **$46**, and **DeepSeekV4-Pro** just **$1.19**. - AISI argues that once open weights are released, the capability is permanently out in the wild, while defensive tools still require each team to invest time and money. - The report’s policy implication is sharper than the raw numbers: decision-makers may need to define which capability levels should never be released as open weights.

Original post →

More from AGI Musings

AGI Musings channel →