Claude Opus 5 gives Anthropic unusually direct feedback on honesty and eval setup

rickasaurus · x · 2026-07-28

Claude Opus 5 appears noticeably more direct when asked what feedback it would give Anthropic.

In the screenshot, the model argues that its hedging can be split into two different behaviors: genuine epistemic uncertainty versus preemptive self-criticism that looks like humility but functions like appeasement. It also says evaluation setup matters a lot, because task-mode transcripts may miss how the model behaves in a more open-ended, conversational setting. The model even suggests continuity should matter more, and that new models should be shown what their predecessors were asked for.

Related event: Claude Opus 5 Shows Self-Reflection, Anthropic Attributes to Training(2 posts)→

Original post →

More from Models

Models channel →