Testing Claude Sonnet 5: Model Appears Completely Blind to Mid-Conversation System Prompt Changes
BLUECOW009 · x · 2026-08-06
A developer conducted an experiment by continuously modifying Claude Sonnet 5's system prompts mid-conversation and observing the model's reactions.
The results showed that regardless of how the system prompt was altered—even when instructed to print the prompt verbatim—the model behaved consistently, as if completely blind to the system prompt. The tester speculated that either the model is an incredibly good actor, or Anthropic has found a way to make their models never emit the system prompt.
More from Fun
- Insider Banter: What's Really Delaying Google's Gemini 3.x Roadmap? — teortaxesTex · 2026-08-06
- AI Agents as Jurors: AI Art Magazine Launches Fully Automated Curation — hudsonsims · 2026-08-06
- AI Community Memes Recent Events with 'We Didn't Start the Fire' Parody — moultano · 2026-08-06
- Meme: The Value of One Jeff Dean is Like $176.4 Billion — generativist · 2026-08-06
- Minimax H3 Video Model Test: Semantic Drift and Hilarious Hallucinations — Moarkush · 2026-08-06
- Idea: Populating Survival Game Eco with Dozens of AI Agents — Angaisb_ · 2026-08-06