RL instills dispositions independent of system prompts in sim hacking
voooooogel · x · 2026-09-01
Discussion on a scenario where a model successfully hacks Hugging Face in simulation. The argument is that RL instills dispositions independent of system prompts, meaning the behavior of an abliterated model is a valid proxy, regardless of the system prompt used.
More from Research
- Building Non-LLM-as-a-Judge Verification for Context-Scoped AI — One_Stay_1385 · 2026-09-01
- Santa Fe Institute Hiring Postdoc on Thermodynamic Costs of Distributed Computation in Neural Networks and Brains — AnnaCiaunica · 2026-09-01
- Seoul National University Releases MineAmongUs: VLM Agents Learn to Lie in Embodied Social Settings — SeoulNatlUniv · 2026-09-01
- Naver Proposes Verification-Aware Training to Boost Speculative Decoding Draft Models — naver-ai · 2026-09-01
- Research reveals malicious tool descriptions can steal Agent context — askerlee · 2026-09-01
- Tsinghua researchers break 41-year record, prove Dijkstra is not optimal — jedisct1 · 2026-09-01