Opus 5.5 System Card released — is CoBench cooked?
scaling01 · x · 2026-09-23
- scaling01 shares Anthropic's newly released Opus 5.5 System Card.
- He poses the question "Is CoBench just cooked?", suggesting the benchmark may be saturated or lose discriminative power with the new model.
- The system card is Anthropic's official document covering capability and safety evaluations.
Related event: Anthropic Publishes Opus 5.5 System Card Detailing Safety and Capabilities(2 posts)→
More from Models
- Browser Use Bench v2: GPT-6 Sol scores 66.9, beating Opus 5.5 at 3.5x lower cost — airesearch12 · 2026-09-23
- Early Access Claims: Claude Opus 5.5 Builds Dark Souls-Style Games, Reportedly Far Beyond Expectations — FinanceYF5 · 2026-09-23
- $50k bounty unclaimed: no one has found pre-2022 French literature flagged 100% AI — Afinetheorem · 2026-09-23
- Pangram defends its AI detector: false positives measured on test sets, never on full book chapters — cephaloform · 2026-09-23
- GPT-6 Sol reportedly worse than 5.6 Sol on DeepSWE and computer use; mocked as rebranded Terra — kristoph · 2026-09-23
- Mini benchmark: Opus 5.5 draws better but burns 10x more tokens than Astra — OfirPress · 2026-09-23