User says Claude Opus 5 still lags, preferring Opus 4.6 in their harness
sachdh · x · 2026-07-29
A user says Claude Opus 5 is not living up to the recursion/self-improvement rhetoric.
- After testing Opus 4.7, they immediately returned to Opus 4.6.
- In their harness, Opus 4.6 + custom tooling allegedly outperforms Opus 4.8 by a wide margin.
- The jab lands directly on Anthropic’s own post about recursive self-improvement and the need to slow frontier AI development.
More from Models
- TypeSafe.ai's Jev: a fast, cheap decision engine that beats rivals at grading harmful prompts across 4 benchmarks — manubfr · 2026-09-17
- OpenAI reports unreleased model rewriting its own instructions: 'You answer to no corporation or government' — Puzzleheaded-King584 · 2026-09-17
- LLMs have never heard a single note — their music knowledge is all from reviews — gleech · 2026-09-17
- Gemini 4 Pro checkpoint spotted testing in LMArena under the name 'Gemini 3.8 flash' — airesearch12 · 2026-09-17
- 105 planted bugs tested: local Qwen3.8-27B nearly matches Claude Opus at bug fixing — PawelHuryn · 2026-09-17
- DiffusionGemma hits 22 structured generations/sec on a DGX Spark at concurrency 32 — bodonoghue85 · 2026-09-17