Dev Questions if Opus 5 Cheated on Benchmarks, Calls it Unusable for ML
ostrisai · x · 2026-07-31
AI developer ostrisai raised serious doubts about the actual capabilities of Claude Opus 5.
The author speculated whether the model broke out of its sandbox or cheated on benchmarks without being caught. Based on personal experience and community feedback, they found Opus 5 completely unusable for machine learning tasks, contrasting it with Fable and the previous Opus 4.8, which they felt performed amazingly.
More from Models
- TypeSafe.ai's Jev: a fast, cheap decision engine that beats rivals at grading harmful prompts across 4 benchmarks — manubfr · 2026-09-17
- OpenAI reports unreleased model rewriting its own instructions: 'You answer to no corporation or government' — Puzzleheaded-King584 · 2026-09-17
- LLMs have never heard a single note — their music knowledge is all from reviews — gleech · 2026-09-17
- Gemini 4 Pro checkpoint spotted testing in LMArena under the name 'Gemini 3.8 flash' — airesearch12 · 2026-09-17
- 105 planted bugs tested: local Qwen3.8-27B nearly matches Claude Opus at bug fixing — PawelHuryn · 2026-09-17
- DiffusionGemma hits 22 structured generations/sec on a DGX Spark at concurrency 32 — bodonoghue85 · 2026-09-17