Both Columns, One Vendor: What Opus 5.5 Benchmarks Can and Cannot Show
maier_ak · x · 2026-10-02
The author links his essay 'Both Columns, One Vendor: What the Opus 5.5 Benchmarks Can and Cannot Show', arguing that hosted model quality can be set administratively by the seller via controls invisible to buyers, and questioning the verifiability of Anthropic's Opus 5.5 launch-page benchmark claims.
Related event: Opus 5.5 Benchmarks Questioned as Vendor Grades Its Own Homework(3 posts)→
More from Models
- Adaptive effort is likely the killer app for dynamic looped transformers — willcb · 2026-10-02
- Dev says Claude and Claude Code are getting worse day by day — Pavan_Belagatti · 2026-10-02
- Abliterated Large V2 lands on Venice: refusal-free AI for red teamers, anonymously — 0xAllen_ · 2026-10-02
- GLM 5.3 lands in Cursor, with GLM 5.3 Max topping CursorBench 4.0 among open-weight models — zainhas · 2026-10-02
- Banned Anthropic users can't get refunds on $200 plans, dev calls for a refund guide — lxfater · 2026-10-02
- Claude Opus 5.5 Tops RSI-Exam at 0.536, Edging Out GPT-6 on 88 Research Tasks — HuaxiuYaoML · 2026-10-02