Interpretability researcher: sandbagging signals from probes would block model deployment

thebasepoint · x · 2026-09-27

In a discussion on interpretability tradeoffs, thebasepoint contrasts linear probes: they interpret internal vectors via probing and steering experiments validated on large corpora, but only yield a scalar readout. He argues that if nonlinear probes indicated models were sandbagging on internal deployments, careful blackbox followup experiments would be run—and if consistent, would block deployment of that model.

Related event: Researchers Debate the Value of NLA and Two Routes of Interpretability(8 posts)→

Original post →

More from Safety

Safety channel →