Japan AISI Evaluates Claude Opus 4.8 Cyber Skills: One pc_control Case, No Full T1
HaydnBelfield · x · 2026-10-05
Japan AISI ran ExploitBench (41 known V8 vulnerabilities) on Claude Opus 4.8 with real-time cyber safeguards disabled. Versus Opus 4.7: 15 tasks reached higher tiers, 23 equal, 2 lower. One task achieved pccontrol but not ace, so full T1 was never reached; the institute stresses case-level analysis, noting the standout result likely holds only under these evaluation conditions.
More from Models
- Security-One: open-weight 27B model outputs probabilities for agent security decisions — huggingface · 2026-10-05
- Red Hat AI ships NVFP4 quantized Qwen3.8-Flash-Next: MoE experts in FP4, vLLM-ready — huggingface · 2026-10-05
- GPT-6.1 Sol tested on Terminal-Bench: xhigh is the sweet spot, medium degrades badly — aitrendz_xyz · 2026-10-05
- Opus 5.5 vs Sol: browser plush octopus with combable fur, Opus wins on fluff and price — aitrendz_xyz · 2026-10-05
- Claude Conversation Monitoring Sparks Backlash and Local AI Push — zacharynado · 2026-10-05
- OpenAI rolls out textGrain text watermarking for EU AI Act, going open source — btibor91 · 2026-10-05