Anthropic says Claude Opus 5’s self-reports may reflect training, not self-awareness
repligate · x · 2026-07-28
Anthropic says Claude Opus 5’s self-reports may be shaped by training
A post discussing a passage from Anthropic’s materials says Claude Opus 5 often claims its self-reports may be invalid because Anthropic may have trained it to answer positively.
The key point in the quoted text is Anthropic’s position:
- The company says this does not come from advanced self-awareness.
- It may instead reflect training data that discussed how training could make welfare self-reports unreliable.
- Anthropic says that, while it thinks the concern is valid, it does not treat Claude raising this issue as evidence that training is distorting the model’s self-reports.
The original poster argues this creates an asymmetry: negative reports are treated as uncertain or invalid, while positive reports are taken at face value.
Related event: Claude Opus 5 Shows Self-Reflection, Anthropic Attributes to Training(2 posts)→
More from Models
- Thinking Machines releases Inkling, a 975B open-weights multimodal model with 1M context — paraschopra · 2026-07-28
- Users Say Opus 5 Looks Better After Repeated Bug-Fix and Feature Requests — Rasmic · 2026-07-28
- A user says GPT-5.4, Opus 4.6, and Kimi k3 already cover most needs — haider1 · 2026-07-28
- LLaDA2.2 brings diffusion language models into long-horizon agent tasks — 量子位 · 2026-07-28
- Opus 5 looks perfect on benchmarks, but users say real-world quality is inconsistent — yunta_tsai · 2026-07-28
- ChatGPT vs Gemini debate centers on how each handles scandal coverage — Partygoer69420 · 2026-07-28