AISI says open models narrow the cyber-range gap
Posts discussing a public analysis from the UK AI Safety Institute (AISI) say the most important takeaway is that open-weight models are closing the gap with closed frontier systems on difficult, long-horizon cybersecurity tasks. In the reported results, the estimated gap is now about 4–7 months, an improvement from the roughly 6–10 month range often referenced for most of 2025. That makes this update notable as a signal that open models are advancing on more realistic cyber-range evaluations, not just short benchmarks.
Key reported results
Several posts point to AISI’s 32-step “The Last Ones” cyber range. According to @daniel_mac8’s summary, GLM-5.2 performed at about the level of Claude Opus 4.5 released roughly seven months earlier. Posts also repeatedly mention DeepSeek V4-Pro as one of the open models highlighted in the latest update, though no exact placement or score is provided in the materials here. Separately, @Outside-Iron-8242 says AISI’s latest results show GPT-5.6 Sol outperforming Mythos 5 on AISI’s cyber challenge.
Additional commentary and limits
A separate post from @AravSrinivas, citing internal evaluation rather than AISI, says Kimi K3 is already top-tier on cybersecurity tasks and argues that claims it is merely “benchmark overfit” do not match its real capability. The same post says Sol is stronger still, but at a clearly higher cost. At the same time, the posts in this cluster do not include AISI’s full leaderboard, raw scores, or DeepSeek V4-Pro’s precise standing, so finer ranking claims still need to be checked against the original AISI release.
2026-07-17 ~ 2026-07-19 · 6 related posts
- UK AISI: Open-Source Model Security Gap Narrows — scaling01 · 2026-07-17
- Comparing Large Models' Practical Cybersecurity Capabilities — daniel_mac8 · 2026-07-17
- Open-Source Models Narrow Cybersecurity Gap — HZoete · 2026-07-18
- AISI Updates Open-Source Model Security Gap — Outside-Iron-8242 · 2026-07-18
- Gap in Long-Range Cybersecurity Capabilities Narrows — rohanpaul_ai · 2026-07-18
- Cybersecurity Evaluation of Open-Weight Models — AravSrinivas · 2026-07-19