Using interpretability probes as privacy-preserving monitors to check models without seeing outputs

anpaure · x · 2026-09-02

A safety researcher floated an idea: if you need a monitor to check whether someone's models are doing bad things, you could build an interpretability-based probe to detect it — with the benefit that the model owner doesn't have to leak their outputs.

This came up while thinking through non-obvious caveats of verification mechanisms; the author works on verification research.

Original post →

More from AGI Musings

AGI Musings channel →