Goodfire on Running Probes During Training
teortaxesTex · x · 2026-07-16
GoodfireAI cautions that experimental setups for interpretability or monitoring probes must be handled with extreme care to avoid invalidating conclusions.
Their solution is to run probes on a frozen copy of the model during training, rather than directly on the active student model. This prevents the student from learning to "evade detection" as training scales.
They also noted that empirically, this method causes no degradation even on post-trained models, and they provided a more detailed discussion article on the topic.
Related event: Goodfire Shares Probe Monitoring Methods in RLFR Training(2 posts)→
More from Research
- Why a 1GW Chinese AI data center may be plausible after all — teortaxesTex · 2026-07-22
- LFM2.5-8B-A1B doubles its tokenizer vocab and cuts on-device decoding time up to 3.7x — maximelabonne · 2026-07-22
- Chinese AI labs are now treating distillation obfuscation as the top research topic — pmddomingos · 2026-07-22
- Structural ensembles beat single predictions in TCR:pMHC generalization study — quaidmorris · 2026-07-22
- RSS launches under OMSF to push structural biology data modeling at scale — MoAlQuraishi · 2026-07-22
- enFoldX tops 8 neoantigen scans and an unseen-peptide benchmark — quaidmorris · 2026-07-22