Goodfire Explains Its RLFR Probe Methodology

burny_tech · x · 2026-07-16

GoodfireAI explained that during RLFR training, they run probes on a frozen copy of the original model rather than the student model itself. This prevents the model from learning to "evade monitoring" as training scales.

They also addressed a common concern: while naive approaches might undermine the effectiveness of interpretability tools, their experimental results show that this setup does not cause any drop in effectiveness, even on post-trained models. The post includes a link to a more detailed explanation.

Related event: Goodfire Shares Probe Monitoring Methods in RLFR Training(2 posts)→

Original post →

More from Safety

Safety channel →