Causal inner product from linear representation hypothesis validated across seven LLMs through 2026
ChenhaoTan · x · 2026-10-09
A follow-up to Park, Choe & Veitch's ICML 2024 paper on the linear representation hypothesis extends the analysis through 2026. The original work proposed a causal inner product — a way of measuring directions so independent concepts become more nearly perpendicular — and showed that moving along a concept direction can flip model predictions. The extension finds the method improves concept separation on all seven models tested, though its advantage over the ordinary inner product shrinks outside the LLaMA family. The paper formalizes linear representation via counterfactuals, connecting it to linear probing and model steering.
More from Research
- Test-Time Structure-Space Search lifts antibody design success from 16% to 78% — anshulkundaje · 2026-10-09
- aimotive driving dataset: train labels come from a tracker that sees the future — RexDouglass · 2026-10-09
- Yale-led study finds symbolic structure inside LLM representations — tallinzen · 2026-10-09
- Discrete Diffusion Meetup at COLM: Oct 8, Grand Ballroom — yuntiandeng · 2026-10-09
- Astra 机器人控制实测:Johns Hopkins 深度评测其运动与控制理解 — jmin__cho · 2026-10-09
- Tavily's 93% live-retrieval agent study challenged: prompts were domain-anchored — edwin · 2026-10-09