'We found a steering vector' is the interpretability paper template that keeps giving, researcher notes
gleech · x · 2026-09-19
Researcher Nina Panickssery calls out a recurring interpretability paper template: "we found a steering vector for <concept> (usually using contrastive pairs)."
She argues it's silly to be surprised by these results — capable LLMs obviously hold abstract representations of every concept appearing in human texts, since that's how they work. Finding steering vectors is a near-inevitable consequence, not a startling discovery.
More from Research
- Cosmos DB VLDB paper shows better scheduling can be worth $100M+ a year — blaizedsouza · 2026-09-19
- Debate: Verification, Not Data, Is the Missing Leap for Useful AI — gerardsans · 2026-09-19
- Tiny tuned classifier beats Jev: GLiNER 2.5 hits 99.7% vs 83.6%, 8.8x faster locally — rickasaurus · 2026-09-19
- Researcher proposes open eval cards and public benchmark repository to fix AI evaluation trust — evijit · 2026-09-19
- How to Benchmark Enterprise AI Memory Beyond LoCoMo — blaizedsouza · 2026-09-19
- MIT's injectable magnetic nanoantennas kill 52% of drug-resistant brain cancer cells — RosalindPicard · 2026-09-19