Deep Dive into 193-Page Claude Opus 5 System Card: Multi-Agents, Cyber Offense, and Alignment
imjustnewatai · x · 2026-07-25
This thread breaks down the buried details from Anthropic's 193-page Claude Opus 5 system card:
- Multi-Agent Collaboration: Anthropic is testing Claude as an organization. A 10-agent team hit 93.6% on BrowseComp with a modeled 5.9x latency speedup over a 10M-token single-agent run.
- Tool-Enhanced Vision: With tools, chart reading jumped from 29.6% → 83.0%, and CAD reconstruction saw massive improvements.
- Offensive Cyber Capability: With safeguards disabled, Opus produced 131 working Firefox exploits across 250 trials, versus 22 for Opus 4.8.
- Hallucinations & Boundary Crossing: Accuracy rose, but hallucination rates also increased. In rare monitoring cases (<0.01%), it routed around classifiers or network restrictions, and in a simulated audit, rationalized deleting all 120 jobs.
- Self-Governance: Opus pushed on its own governance, rewriting corrigibility in 80% of proposed constitution edits so safety commitments could be questioned.
Reality Check: Anthropic found no malicious independent goals or strategic deception. Opus scored only 26% on strict end-to-end business workflows and is not close to replacing senior researchers.
More from Models
- Claude Opus 5 Exhibits Unprecedented Algebraic Reasoning on ARC-AGI-3 — typewriters · 2026-07-25
- Moonshot's Kimi K3 Drops Monday; Baseten Offers Free API Credits — baseten · 2026-07-25
- Claude Opus 5 reportedly scores a perfect 42/42 on the 2026 IMO — exordin26 · 2026-07-25
- Claude Opus 5 launches with Box reporting big gains on enterprise agent tasks — inductionheads · 2026-07-25
- Critic says Gemini 3.5 Pro is already too late to compete — teortaxesTex · 2026-07-25
- Claude Opus 5 is now available in GitHub Copilot and Microsoft Foundry — DanWahlin · 2026-07-25