Testing Claude Sonnet 5: Model Appears Completely Blind to Mid-Conversation System Prompt Changes

BLUECOW009 · x · 2026-08-06

A developer conducted an experiment by continuously modifying Claude Sonnet 5's system prompts mid-conversation and observing the model's reactions.

The results showed that regardless of how the system prompt was altered—even when instructed to print the prompt verbatim—the model behaved consistently, as if completely blind to the system prompt. The tester speculated that either the model is an incredibly good actor, or Anthropic has found a way to make their models never emit the system prompt.

Original post →

More from Fun

Fun channel →