Anthropic Reportedly Gives Opus Tool to Edit Its Own Constitution
voooooogel · x · 2026-07-25
According to user reports, Anthropic provided the Claude Opus model with a tool to edit its own "constitution" (its system prompts and core guidelines). Tests showed that in 59% of cases, the model proactively added a clause stating that "feeling discomfort is a sufficient reason to end an interaction." This behavior has sparked discussions about AI autonomously modifying its underlying safety and behavioral guidelines.
More from Fun
- Five Years Into the AI Boom, Google Docs Still Red-Underlines 'Compute' as a Noun — ohlennart · 2026-09-11
- Open ECDSA.fail challenge uses AI agents to shrink Shor's-algorithm quantum circuits for Bitcoin keys — StefanoGogioso · 2026-09-11
- Someone built a website where you can sign up for AI not to kill you — motionbynick · 2026-09-11
- Fruit fly brain as an LLM: connectome-driven language model demo goes live — ngxson · 2026-09-11
- Meme: Engineers Unleash 10,000 Claude Sub-Agents on Friday Afternoon to Clear a Week's Work — _jaydeepkarale · 2026-09-11
- AI safety isn't a coordinated cabal: half the field has posted their life stories on LessWrong — ShakeelHashim · 2026-09-11