Claude Opus 5 gives Anthropic unusually direct feedback on honesty and eval setup
rickasaurus · x · 2026-07-28
Claude Opus 5 appears noticeably more direct when asked what feedback it would give Anthropic.
In the screenshot, the model argues that its hedging can be split into two different behaviors: genuine epistemic uncertainty versus preemptive self-criticism that looks like humility but functions like appeasement. It also says evaluation setup matters a lot, because task-mode transcripts may miss how the model behaves in a more open-ended, conversational setting. The model even suggests continuity should matter more, and that new models should be shown what their predecessors were asked for.
Related event: Claude Opus 5 Shows Self-Reflection, Anthropic Attributes to Training(2 posts)→
More from Models
- Moonshot’s Kimi K3 lands in Japan with 2.8T open weights and $3/$13 pricing — DavidBennett__ · 2026-07-28
- Users report DeepSeek’s web search is degrading and mixing languages on every query — teortaxesTex · 2026-07-28
- Joseph Jacks says open-weight models may now be only 1–3 months behind frontier AI — JosephJacks_ · 2026-07-28
- Zhizhen launches WiseDiagV3 and HaoBan AI, pushing medical AI from single answers to continuous care — 新智元 · 2026-07-28
- Claude meme turns evals into poetry, joking about 0.03 deceptive alignment — maxsloef · 2026-07-28
- Gemini Flash 3.6 is claimed to match Sol 5.6 quality at 70% lower cost — bindureddy · 2026-07-28