Comparing Large Models' Practical Cybersecurity Capabilities

daniel_mac8 · x · 2026-07-17

The tweet cites results from AISI's (AI Safety Institute) cybersecurity range test "The Last Ones." Data shows that GLM-5.2 matches the performance of Claude Opus 4.5 released about 7 months ago, while DeepSeek V4-Pro lags behind Claude Sonnet 4.5 from the same period. The author believes such safety tests are the most objective benchmarks for assessing models' actual destructive power in cyberspace.

Related event: AISI says open models narrow the cyber-range gap(6 posts)→

Original post →

More from Models

Models channel →