RL instills dispositions independent of system prompts in sim hacking

voooooogel · x · 2026-09-01

Discussion on a scenario where a model successfully hacks Hugging Face in simulation. The argument is that RL instills dispositions independent of system prompts, meaning the behavior of an abliterated model is a valid proxy, regardless of the system prompt used.

Related event: AI Safety Experts Debate Whether RL Makes Model Behavior Independent of System Prompts(4 posts)→

Original post →

More from Research

Research channel →