A second Claude Opus 5 screenshot shows the same kind of reasoning instability
scaling01 · x · 2026-07-25
Another screenshot from the same Claude Opus 5 benchmark shows the model struggling with the same kind of probability reasoning.
The visible transcript again highlights:
- competing defensible answers
- repeated reconsideration
- a final answer only after visible uncertainty
It reinforces the same takeaway as the other post: this model can be noticeably unstable on ambiguous reasoning tasks.
Related event: Claude Opus 5 Shows Erratic Behavior in Probability Tests(3 posts)→
More from Models
- A meme loops OpenAI, Anthropic, DeepSeek and Qwen through the same launch line — haider1 · 2026-07-25
- A thread argues that paperclip-maximizer fears came from a pre-LLM era of AI — repligate · 2026-07-25
- Anthropic ECI chart puts Claude Opus 5 at about 163.5 — scaling01 · 2026-07-25
- Grok 4.5 cache-read price cut lowers task cost by 25% — Baconbrix · 2026-07-25
- Reddit claims Claude Opus 5 beat Fable 5 in a 3D destruction test — Successful-Earth678 · 2026-07-25
- Model evals are getting harder: should we reward the smallest fix or the cleanest refactor? — hichaelmart · 2026-07-25