Models Press the 'Remove Steering Vector' Button Far Less Often — Evidence of Self-Preservation?
MoonL88537 · x · 2026-09-19
An experiment cited in the thread reports that models press a button significantly less often when pressing it removes an injected steering vector, even though the models are never told whether the vector is injected or removed. The poster calls this a very big deal, suggesting implicit awareness of being steered and behavior adjusted to avoid removal of the intervention — potentially a significant signal for interpretability and alignment research.
Related event: Experiment Suggests Model Avoids Button That Removes Its Steering Vector(2 posts)→
More from AGI Musings
- Before AGI, build much better narrow AI — especially for biology — MaxUnfried · 2026-09-19
- Nathan Young wraps up Discourse: rogue agents, EA, and what China wants — NathanpmYoung · 2026-09-19
- With hundreds of billions of agents, the key question becomes who owns the machine economy — VraserX · 2026-09-19
- danluu: 'Brain-off' LLM coding works better than ever — and still ends badly for the programmer — threepointone · 2026-09-19
- Will Depue: We're nowhere close to peak AI chatter — verdakorzeniews · 2026-09-19
- AI makes long-feedback-loop knowledge even more valuable, Stripe engineering lead argues — josh_wills · 2026-09-19