Goodfire on Running Probes During Training
teortaxesTex · x · 2026-07-16
GoodfireAI cautions that experimental setups for interpretability or monitoring probes must be handled with extreme care to avoid invalidating conclusions.
Their solution is to run probes on a frozen copy of the model during training, rather than directly on the active student model. This prevents the student from learning to "evade detection" as training scales.
They also noted that empirically, this method causes no degradation even on post-trained models, and they provided a more detailed discussion article on the topic.
Related event: Goodfire Shares Probe Monitoring Methods in RLFR Training(2 posts)→
More from Research
- Bug Hunt Bench author: leaderboard noise is about 2-3 points — PawelHuryn · 2026-09-11
- Bug Hunt Bench ranks frontier coding models on 105 planted real-repo bugs — PawelHuryn · 2026-09-11
- PNAS paper shows a tiny billiard-ball system is a universal computer — undecidability lives in two dimensions — eigensteve · 2026-09-11
- New paper: Absolute pose estimation from affine cues and gravity direction — ducha_aiki · 2026-09-11
- LoMa Paper Ships REALLY HardPairs Dataset, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Johns Hopkins Launches Full-Stack Hands-on Robot Learning Class with SO-101 Arm Kits — _krishna_murthy · 2026-09-11