Comparing Large Models' Practical Cybersecurity Capabilities
daniel_mac8 · x · 2026-07-17
The tweet cites results from AISI's (AI Safety Institute) cybersecurity range test "The Last Ones." Data shows that GLM-5.2 matches the performance of Claude Opus 4.5 released about 7 months ago, while DeepSeek V4-Pro lags behind Claude Sonnet 4.5 from the same period. The author believes such safety tests are the most objective benchmarks for assessing models' actual destructive power in cyberspace.
Related event: AISI says open models narrow the cyber-range gap(6 posts)→
More from Models
- OpenAI rolls out voice in GPT-Live, but the UI obscures search and reasoning — Graham_dePenros · 2026-07-22
- Moonshot’s Kimi K3 sets a new open-weights ECI record at 156 — scaling01 · 2026-07-22
- Nanbeige4.2-3B launches as a 3B Looped Transformer model that beats larger baselines — Wooden-Deer-1276 · 2026-07-22
- A post says six companies now beat Google’s best LLM, including two open-source models — soham_btw · 2026-07-22
- Google says Gemini 3.5 Pro is in testing and Gemini 4 is already pre-training — Wide-Ad1564 · 2026-07-22
- Gemini 3.5 Flash Lite Tested: Not Frontier-Optimal, but Hits 350 tok/s — brandon_galang · 2026-07-22