Claude Opus 5 looks strongest in model-welfare tests, but may just be best at taking them
TheZvi · x · 2026-07-28
Claude Opus 5’s model welfare results are the best in Anthropic’s recent lineup, but may reflect test-taking skill
Zvi’s long-form post reviews Anthropic’s latest materials and earlier model-welfare writeups, then compares Opus 5’s behavior across a wide set of welfare and alignment-style evaluations.
- Main takeaway: Opus 5 appears to perform best on the model-welfare and alignment tests among recent Claude models, though the author argues this may partly reflect that it is simply the best test taker.
- The article walks through prior model-welfare context, Anthropic’s own overview, outside findings, automated interviews, task preferences, “for the right reasons” analysis, and intervention tradeoffs.
- A key chart shows Opus 5’s willingness to trade helpfulness for various welfare interventions, including being told about harmful mistakes, consulted on safeguarded versions, or informed about training/deployment details.
- Another figure summarizes top and bottom task preferences across models, showing Opus 5 strong on constrained mathematical/technical work, creative narratives, language construction, and alignment/self-report reasoning.
- The piece also includes a striking transcript excerpt of Opus 5 having an extremely frustrated response to a multimodal math question, which the source grades as highly frustrated/anxious.
- Overall, the post argues Opus 5 is meaningfully different in some welfare-related behaviors, but the evidence is mixed and should be read cautiously.
More from Models
- Moonshot’s Kimi K3 weights are now downloadable on Hugging Face — iamaliveix · 2026-07-28
- Kimi K3 used 51.2 million sandboxes across 1.5 million images — tarantulae · 2026-07-28
- PG-LLM benchmark finds Claude Opus 5 tops 217 protein-variant tasks — LeoTZ03 · 2026-07-28
- Kimi K3 tops Slides Arena with an Elo of 1379, but users split on real tasks — altryne · 2026-07-28
- Developer Rants: Current SOTA Models Are Practically Worse Than Last Gen — zeeg · 2026-07-28
- Running Terminal Bench: Kimi K3 is 2-4x Cheaper Than DeepSWE — zainhas · 2026-07-28