Zvi's System Card Read: Claude Opus 5.5 Tops Benchmarks, Bio Evals Drop Helpful-Only Testing
Don't Worry About the Vase (Zvi) · rss · 2026-09-24
Zvi Mowshowers offers a section-by-section read of Anthropic's Claude Opus 5.5 system card. The model is billed as the strongest by standard benchmarks while being cheaper than Opus 5, with Anthropic claiming it outright matches or beats Fable 5.1.
Key takeaways:
- Classifiers and deployment: Five risk-classifier areas; cyber misuse falls back to Opus 4.8, chemical/bio falls back to Opus 5; distillation attacks are blocked with no fallback.
- RSP: Rated CB-1 but not CB-2; autonomy risk on-trend and well below the Autonomy-2 threshold, deployed with safeguards comparable to prior releases.
- Major policy shift: Anthropic will no longer test helpful-only Claude versions, switching to refusal-free eval designs (a 16-hour antibacterial-treatment red-team exercise, black-box RNA sequence design, AAV capsid packaging tasks), arguing relevant capabilities are inherently dual-use. Zvi notes real costs: some evals were dropped and teams lost time to refusals.
- Bio performance: Opus 5.5 sets a new high on black-box RNA sequence design, but still struggles with open-ended scientific reasoning—overrelying on paper abstracts and even designing DNA that doesn't encode the intended protein. In red-teaming, experts raised the average but the top team was generalists.
- AI R&D: Automated research capability is improving but only on-trend; epistemic quality and instruction following remain the main bottlenecks.
Zvi's overall verdict: the situation hasn't importantly changed, though he doubts the measured upper bound on bio risk reflects what a well-resourced team could actually extract.
More from Models
- New Sol wins users over: same capabilities as Astra, faster and cheaper — max_paperclips · 2026-09-24
- goodside Asks Claude Opus 5.5 to Make a Poignant 1-Minute Video About Its Life — goodside · 2026-09-24
- Dev: Claude 5.5 Is So Good I Didn't Expect to Become a Claude Shill Again — haydendevs · 2026-09-24
- Together releases Tev1-4B-experimental weights to show how easy it is to train decision models — togethercompute · 2026-09-24
- Mathematicians report unlimited tokens; likely just OpenAI's lenient Codex resets — JacquesThibs · 2026-09-24
- Train your own Jev-like classifier for $17 with Qwen3.5 4B fine-tuning — nutlope · 2026-09-24