Users Accuse Anthropic of Cooked Evals, Claiming Real API Performance Lags
GabGarrett · x · 2026-08-12
A user criticized Anthropic on social media, claiming that since version 4.7, the company's published evaluation results have been "cooked" and do not reflect what is actually being served to customers.
The user further alleged that since version 4.8, Anthropic has outright refused any benchmarks on the public API, a move that intensifies skepticism regarding the model's true capabilities and real-world performance.
More from Models
- US Treasury Secretary Bessent Endorses Open-Source AI as a Win for Innovation — max_paperclips · 2026-08-12
- TinyTitle: An Ultra-Lightweight Chat Title Model Running in Under 5MB of RAM — H-L_echelle · 2026-08-12
- Anthropic's Watermark Strategy Flawed: Could Become Top Distillation Target — cocktailpeanut · 2026-08-12
- Research Reveals the Personality Evolution of the Grok Model Family — DevDminGod · 2026-08-12
- Rumor: Sonnet 5 Price Hike Delayed to Offset Opus 5 Backlash — creblohulk · 2026-08-12
- User Slams OpenAI's Safety Filters While Auditing Insulin Pump — max_paperclips · 2026-08-12