Opus 5.5 Benchmarks Can't Be Verified Outside Anthropic — Author Proposes Cheap Independent Audit
maier_ak · x · 2026-10-02
Andreas Maier's essay argues hosted model quality can be set administratively by the seller via controls buyers can't observe. He traces Anthropic's release cadence: Opus 4.7 (Apr 16) drew a record 35 unwarranted-refusal reports in a month, 4.8 followed 42 days later; Opus 5 (Jul 24) drew complaints of verbose answers and massive unnecessary rewrites, with 5.5 arriving 60 days later. His point isn't fraud but that the launch page answers questions nobody outside the company can check. He proposes a cheap audit: pinned snapshots, stated effort per benchmark column, and a held-out suite run daily by someone not selling the model.
Key points
- All launch-page numbers are self-measured, third-party irreproducible
- 5.5's price cuts (input $5→$4, output $25→$20, cache reads -60%) undercut the 'cripple the old' theory
- But the claimed 40% serving-cost reduction implies savings come from curtailed deliberation
- Calls for an industry-wide cheap independent verification mechanism
Related event: Opus 5.5 Benchmarks Questioned as Vendor Grades Its Own Homework(3 posts)→
More from Models
- Latest AI models are so good at reading binaries that 'every game will be effectively open source' — bradneuberg · 2026-10-02
- Ajeya Cotra's 2026 AI forecast review: 'plan a birthday party' has already fallen — ajeya_cotra · 2026-10-02
- Pi 1.0 and Pi Durable Hit HN as Gemini 4 Argon, GPT-6.1 Sol and FLUX 3 Land — Latent Space · 2026-10-02
- Cantina releases apex-flash-1, an open-weights security model post-trained on real paid vulnerabilities — ctjlewis · 2026-10-02
- Unverified uncensored GLM-5.3 EXL3 3.0bpw quantized upload appears on Hugging Face — kevinnbass · 2026-10-02
- GPT-6.1 Sol and Claude Sonnet 5.5 Roll Out to Copilot, ChatGPT Debut Hinted — koltregaskes · 2026-10-02