'We found a steering vector' is the interpretability paper template that keeps giving, researcher notes

gleech · x · 2026-09-19

Researcher Nina Panickssery calls out a recurring interpretability paper template: "we found a steering vector for <concept> (usually using contrastive pairs)."

She argues it's silly to be surprised by these results — capable LLMs obviously hold abstract representations of every concept appearing in human texts, since that's how they work. Finding steering vectors is a near-inevitable consequence, not a startling discovery.

Original post →

More from Research

Research channel →