BlueBench Cyber Benchmark: GPT-5.6 Sol Leads, Kimi K3 Tops Open-Weights
vijaybolina · x · 2026-08-16
BlueBench-Intrusion-003 benchmark released, testing AI cyber incident response using real AWS intrusion data. Models investigate CloudTrail, GuardDuty, and S3 logs.
Key results:
- GPT-5.6 Sol leads at 88.3%; GPT family dominates the pareto frontier.
- Kimi K3 leads open-weight models at 85.1%.
- Anthropic uniquely affected by service-level cybersecurity refusals.
More from Models
- dots3-note-prev model released: SWE-bench 78.4, designed for local use — victormustar · 2026-08-16
- llama.cpp integrates Dots3 Note model, scoring 78.4 on SWE-bench Verified — victormustar · 2026-08-16
- Post-training boosts GLM-5.3 to rival larger models, showing size isn't the bottleneck — haider1 · 2026-08-16
- Commits suggest Qwen 35B model removed, likely not releasing — Local-Cardiologist-5 · 2026-08-16
- Blind Test: Anime Girl 3D Scene Generation Across Models — Jeanodel · 2026-08-16
- 1-bit Quantized Qwen 27B Runs on 12GB VRAM at 92 tokens/s — zyxciss · 2026-08-16