Goodfire on Running Probes During Training
teortaxesTex · x · 2026-07-16
GoodfireAI cautions that experimental setups for interpretability or monitoring probes must be handled with extreme care to avoid invalidating conclusions.
Their solution is to run probes on a frozen copy of the model during training, rather than directly on the active student model. This prevents the student from learning to "evade detection" as training scales.
They also noted that empirically, this method causes no degradation even on post-trained models, and they provided a more detailed discussion article on the topic.
Related event: Goodfire Shares Probe Monitoring Methods in RLFR Training(2 posts)→
More from Research
- Causal-only attention for non-generative tasks is wasteful, argues HF engineer — antoine_chaffin · 2026-09-11
- Catholic University of Chile researcher: scaling AI feedback is key to sustainable medical education — julianvarascom · 2026-09-11
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11
- SignNet 1M Dataset Released for Sign Language Research — ducha_aiki · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- InFlux++ Method Released — ducha_aiki · 2026-09-11