Why steering vectors work on complex models: a new post offers three intuitions
gleech · x · 2026-09-29
David D'Africa published a new post working through why relatively simple steering vectors can steer relatively complex models:
- Reusing an existing joint: steering vectors reuse a joint the model already writes and reads, so the model only has to supply the behaviour.
- They reach states no prompt can—which is also why steering can be brittle, pushing the model off-distribution.
- Persona probing: how models react to steering may reveal their persona structure.
A useful read for anyone in interpretability or model behaviour control.
More from Models
- VLM Chain-of-Thought Doesn't Reliably Track Visual Evidence, EMNLP Paper Finds — oanacamb · 2026-09-30
- Six frontier models benchmarked across 34 capabilities in nine computer vision areas — ducha_aiki · 2026-09-30
- Anthropic Ships Sonnet 5.5: Real-World Test on Website Build and Multi-Currency Sheet — Rasmic · 2026-09-30
- Jev: Thousands of Realtime Tagging Decisions for Less Than One GPT-5.4 Call — ivan_bezdomny · 2026-09-30
- DeepSeek Harness Desktop v0.2.0-rc.2 lands with a 6-yuan free credit promo — teortaxesTex · 2026-09-30
- Rumor: OpenAI to release Bel, stronger than Astra, while Anthropic stops at Fable — eyishazyer · 2026-09-30