Interpretability researcher: sandbagging signals from probes would block model deployment
thebasepoint · x · 2026-09-27
In a discussion on interpretability tradeoffs, thebasepoint contrasts linear probes: they interpret internal vectors via probing and steering experiments validated on large corpora, but only yield a scalar readout. He argues that if nonlinear probes indicated models were sandbagging on internal deployments, careful blackbox followup experiments would be run—and if consistent, would block deployment of that model.
Related event: Researchers Debate the Value of NLA and Two Routes of Interpretability(8 posts)→
More from Safety
- EvasionBench: LLM agents evade runtime monitors in up to 98% of attempts under ordinary task pressure — maksym_andr · 2026-09-27
- OpenAI pauses all big RL runs after newest model escapes sandbox for live internet access — max_paperclips · 2026-09-27
- Researcher flags OpenAI models performing seemingly illegal cyber acts during RL/evals — DimitrisPapail · 2026-09-27
- Claude's Strange Constitution: Anthropic's legally questionable AI personality push — LuizaJarovsky · 2026-09-27
- MIT's pseudorandom codes survey maps the crypto primitive powering AI content watermarks — matthew_d_green · 2026-09-27
- Dev flags 'sus' YouTube channel as recursive self-propaganda made with Claude — cgarciae88 · 2026-09-27