Causal inner product from linear representation hypothesis validated across seven LLMs through 2026

ChenhaoTan · x · 2026-10-09

A follow-up to Park, Choe & Veitch's ICML 2024 paper on the linear representation hypothesis extends the analysis through 2026. The original work proposed a causal inner product — a way of measuring directions so independent concepts become more nearly perpendicular — and showed that moving along a concept direction can flip model predictions. The extension finds the method improves concept separation on all seven models tested, though its advantage over the ordinary inner product shrinks outside the LLaMA family. The paper formalizes linear representation via counterfactuals, connecting it to linear probing and model steering.

Original post →

More from Research

Research channel →