Steering vector intuitions: persona, concept, belief and task vectors are all the same thing

voooooogel · x · 2026-10-02

In a technical thread responding to the claim that "steering vectors are behavioral biases," the author shares useful intuitions: persona, concept, belief and task vectors are all the same construct, just extracted (mean differences, PCA, NLA reconstruction, SAE) and applied (activation addition, patching) differently; the LLM residual stream is a noisy bag of representations, and steering-vector extraction is about distilling it into clean, reusable, interpretable directions; and mean-difference works by cancelling distractor features on either side of the subtraction — e.g. ignoring "user age" features when extracting a model-distress vector.

Original post →

More from Research

Research channel →