One Word Apart: Anthropic Tweaked Its Anti-Profanity System Prompt Between Opus 4.1 and 4.5
steipete · x · 2026-10-01
A comparison of Anthropic's system prompts shows a subtle wording change to the no-profanity clause: Opus 4.1 read "Claude never curses unless the human asks for it," while Opus 4.5 reads "Claude never curses unless the person asks Claude to curse." The tweak highlights how delicate prompt language is — specifying what the request must ask for may matter for how the constraint is interpreted in practice.
More from Fun
- AI speeds up research so much you can answer an arXiv question the same day — KyleCranmer · 2026-10-01
- User reports being unable to cancel ChatGPT subscription — imjustnewatai · 2026-10-01
- Polish netizen jokes Poland was ahead on AI: it's always been 'SI' in Polish — tlakomy · 2026-10-01
- OpenAI researcher jokes model release finally excuses his late replies to friends and family — divy93t · 2026-10-01
- Skip the AI "muse" — this user wants a scheming, silver-tongued vizier persona — tedmitew · 2026-10-01
- 'AGI achieved boys' — Reddit celebrates with a meme — GeneReddit123 · 2026-10-01