Opus 5.5 Benchmarks Can't Be Verified Outside Anthropic — Author Proposes Cheap Independent Audit

maier_ak · x · 2026-10-02

Andreas Maier's essay argues hosted model quality can be set administratively by the seller via controls buyers can't observe. He traces Anthropic's release cadence: Opus 4.7 (Apr 16) drew a record 35 unwarranted-refusal reports in a month, 4.8 followed 42 days later; Opus 5 (Jul 24) drew complaints of verbose answers and massive unnecessary rewrites, with 5.5 arriving 60 days later. His point isn't fraud but that the launch page answers questions nobody outside the company can check. He proposes a cheap audit: pinned snapshots, stated effort per benchmark column, and a held-out suite run daily by someone not selling the model.

Key points

Related event: Opus 5.5 Benchmarks Questioned as Vendor Grades Its Own Homework(3 posts)→

Original post →

More from Models

Models channel →