Dev Slams Claude 3.5 Opus as 'Benchmaxxed', Claims It Feels Dumber in Practice

evilsocket · x · 2026-07-31

Prominent developer evilsocket expressed deep frustration with the Claude 3.5 Opus model on social media. He complained that during his daily, hands-on testing, the model feels and acts significantly dumber than version 4.8 (likely referring to GPT--4.8 or a similar competitor).

He suspects the model was heavily 'benchmaxxed'—optimized specifically to score high on public benchmarks—at the expense of real-world performance and usability in complex workflows.

Related event: Claude Opus Series Accused of Degraded Experience: Laziness and Amnesia Spark Trust Crisis(14 posts)→

Original post →

More from Models

Models channel →