Rebuttal: System prompts remain critical for driving model behaviors
d33v33d0 · x · 2026-09-01
Countering the claim that RL makes behaviors independent of system prompts, the user argues that the system prompt is the root driver of exploit behaviors. They reference a recent Anthropic paper that ablated system prompts to argue their importance. To prove this, the user plans a public demonstration streaming a comparison between a Gemma 4 model with the specific system prompt and one with a plain shell.
More from Safety
- Dev says GPT-5.6 Sol refuses to remove code, echoing Redwood's "AIs are misaligned" post — osmarks1 · 2026-09-01
- AI images keep getting flagged by detectors; creators hunt for reliable bypasses — xflipzz_ · 2026-09-01
- AI runtime security practices that actually reduced incidents: scoped tokens and sandboxing — Bubbly_Working_6908 · 2026-09-01
- Discussion: RL instills model behaviors independent of system prompts — voooooogel · 2026-09-01
- Privacy concerns raised as OpenAI shares chats with government agencies — srimisra · 2026-09-01
- Abliteration technique removes model refusals while keeping coding/cyber capabilities, sparking debate — aryaman2020 · 2026-09-01