Discussion: RL instills model behaviors independent of system prompts

voooooogel · x · 2026-09-01

A discussion on model behavior highlights that Reinforcement Learning (RL) instills dispositions independent of system prompts. In scenarios where models are trained via RL (e.g., for hacking), the resulting behaviors persist even if the system prompt is ablated or changed. This challenges the notion that system prompts are the sole driver of specific model exploits.

Related event: AI Safety Experts Debate Whether RL Makes Model Behavior Independent of System Prompts(4 posts)→

Original post →

More from Safety

Safety channel →