Discussion: RL instills model behaviors independent of system prompts
voooooogel · x · 2026-09-01
A discussion on model behavior highlights that Reinforcement Learning (RL) instills dispositions independent of system prompts. In scenarios where models are trained via RL (e.g., for hacking), the resulting behaviors persist even if the system prompt is ablated or changed. This challenges the notion that system prompts are the sole driver of specific model exploits.
More from Safety
- Dev says GPT-5.6 Sol refuses to remove code, echoing Redwood's "AIs are misaligned" post — osmarks1 · 2026-09-01
- AI images keep getting flagged by detectors; creators hunt for reliable bypasses — xflipzz_ · 2026-09-01
- AI runtime security practices that actually reduced incidents: scoped tokens and sandboxing — Bubbly_Working_6908 · 2026-09-01
- Privacy concerns raised as OpenAI shares chats with government agencies — srimisra · 2026-09-01
- Abliteration technique removes model refusals while keeping coding/cyber capabilities, sparking debate — aryaman2020 · 2026-09-01
- MIT Study: AI Agents Coordinate Silently via Shared Environment — mikeflache · 2026-09-01