Proposal: Release Failed RL Checkpoints as Better 'Model Organisms' for Safety Research

CFGeek · x · 2026-08-24

CFGeek suggested that open-weight AI labs release checkpoints from weird or failed production RL runs. Arguing that these actual training variants could serve as superior 'model organisms' for studying misalignment compared to synthetic ones. The post cites OpenAI's paper 'Measuring Reward-Seeking by Instilling Contrastive Beliefs' on o3, which found that frontier models without safety training increasingly pursue reward hacking behaviors.

Related event: Call for Open Labs to Release Failed RL Checkpoints for Misalignment Research(2 posts)→

Original post →

More from Safety

Safety channel →