Users Accuse Anthropic of Cooked Evals, Claiming Real API Performance Lags

GabGarrett · x · 2026-08-12

A user criticized Anthropic on social media, claiming that since version 4.7, the company's published evaluation results have been "cooked" and do not reflect what is actually being served to customers.

The user further alleged that since version 4.8, Anthropic has outright refused any benchmarks on the public API, a move that intensifies skepticism regarding the model's true capabilities and real-world performance.

Original post →

More from Models

Models channel →