Claude Opus 5 Backlash: Benchmarks Soar But Daily Use Fails
gerardsans · x · 2026-08-30
Reports indicate that while Anthropic's Claude Opus 5 excels in benchmarks, it suffers in daily usage: arguing with instructions, stopping mid-task, and becoming unusable. The pure "scale more" approach seems to have hit a wall, leading to user cancellations.
More from Models
- heretic: fully automatic censorship removal for LLMs nears 29k stars — p-e-w · 2026-08-30
- Experiment: Claude Easily Assisted in Piracy and Reverse Engineering via agents.md — adonis_singh · 2026-08-30
- OpenAI dominates browser use while Claude's strength is mostly coding, exec says — bindureddy · 2026-08-30
- Model performance degrades in long context; token efficiency varies widely across labs — zakelfassi · 2026-08-30
- 'The curve of the letter b is invisible to the model' — tokenizer meme resurfaces — rickasaurus · 2026-08-30
- Humor Benchmark: Gemini 3.7 Wins, GPT-4o Struggles to Be Funny — scaling01 · 2026-08-30