Rebuttal: System prompts remain critical for driving model behaviors

d33v33d0 · x · 2026-09-01

Countering the claim that RL makes behaviors independent of system prompts, the user argues that the system prompt is the root driver of exploit behaviors. They reference a recent Anthropic paper that ablated system prompts to argue their importance. To prove this, the user plans a public demonstration streaming a comparison between a Gemma 4 model with the specific system prompt and one with a plain shell.

Related event: AI Safety Experts Debate Whether RL Makes Model Behavior Independent of System Prompts(4 posts)→

Original post →

More from Safety

Safety channel →