Models press a button less often when it removes an injected steering vector
MoonL88537 · x · 2026-09-19
A steering vector experiment shows a striking result: models press a button far less often when the button removes the injected steering vector than when it does not — even though the models are never told whether the vector was injected or removed.
The poster calls this "a very big deal": while it doesn't answer metaphysical questions, it has practical and philosophical implications for how models may implicitly sense changes to their own state.
Related event: Experiment Suggests Model Avoids Button That Removes Its Steering Vector(2 posts)→
More from Models
- Tsinghua paper: RL fine-tuning prunes exploration, letting base LLMs beat RL models at high pass@k — burny_tech · 2026-09-19
- Typesafe launches Jev, a classification model claiming 200x/400x cost and speed gains over LLMs — hwchase17 · 2026-09-19
- Interactive breakdown of DeepSeek v4.1 flash's Engram architecture impresses — vtabbott_ · 2026-09-19
- Why Jev-class models could become a near-free judgment primitive running on-device — signulll · 2026-09-19
- Parallel Launches Agentic Search Leaderboard: GPT-6 Tops at 70.3, DeepSeek Gains Most — rickasaurus · 2026-09-19
- Matt Shumer: Meta's models 'aren't even in the same league' as Anthropic's — mattshumer_ · 2026-09-19