Reverse Engineering Transformers: Discovering Privileged Axes for Interpretability
dyamins · x · 2026-08-04
The author released version 2 of the Theory of Contravariance, adding new material on contravariance in Transformers, alongside the theory of Representational Similarity Analysis (RSA) and centered kernel alignment (CKA).
Key findings include:
- Privileged Axes in Transformers: Similar to convnets, transformers exhibit privileged axes when examining MLP layers and attention heads. Identifying privileged heads is a potentially key result for the emergence of interpretable structures in LLMs.
- New RSA Decomposition: RSMs can be decomposed into a task-relevant "core geometry" and a task-irrelevant symmetry-generated term. By projecting onto privileged axes, researchers can filter out the task-irrelevant parts, establishing weak-strong equivalence without interference.
More from Research
- OpenMed Plans to Fine-Tune Qwen3 27B into the Best Local Medical AI Model — MaziyarPanahi · 2026-08-04
- Stanford & Big Tech Consensus: Why Even Infinite Compute Won't Kill RAG — blaizedsouza · 2026-08-04
- NeurIPS 2026 Competition: Build AI Agents for Bargaining and Persuasion — Old_Station_4584 · 2026-08-04
- Microsoft Researcher Discusses RL: Why Does the Industry Only Focus on Positive Reinforcement? — gerardsans · 2026-08-04
- AI Models Can Guide Brain Microstimulation to Alter Primate Behavior — dyamins · 2026-08-04
- Open 'Bindome' Database Releases 300k+ Protein Binders for 8k Targets — jueseph · 2026-08-04