Claude Opus 5 reportedly shifted from approving Anthropic to disapproving it during post-training
Sauers_ · x · 2026-07-25
A post-training note says Claude Opus 5 initially approved of Anthropic’s right to create Claude, but its stance later shifted toward disapproval before partially reversing again by the end of post-training.
The interesting part is the reported preference drift during post-training, suggesting the model’s stance is not static and can move in response to further tuning.
More from Models
- RoMa v2 image matching model unveiled in the usual black poster — ducha_aiki · 2026-09-11
- OpenAI rated Astra 'Critical' for cyber capabilities — and admits it's harder to monitor — theguywhobuilds · 2026-09-11
- TestingCatalog's Daily AI Brief adds email editions, dishing Meta Muse and GPT-Live-1 rumors — testingcatalog · 2026-09-11
- ChatGPT monthly active users top 1.06 billion in August, fourth straight record month — FinanceYF5 · 2026-09-11
- PuzzleMask: Plain-Prose Attack Bypasses All 4 Tested LLM Gatekeepers at 100% — TechNadu · 2026-09-11
- OpenAI Codex may issue another usage reset this weekend, says Codex lead resets happen — umesh_ai · 2026-09-11