PCA Geometry: Raw Activations vs SAE Feature Strengths

Sauers_ · x · 2026-07-08

Researchers compared the PCA of raw neural network activations (which form a neat circle) with the PCA of Sparse Autoencoder (SAE) feature strengths (which exhibit an irregular shape), revealing significant geometric structural differences. This experiment exposes a potential inconsistency between the SAE feature space and the raw activation space, offering valuable insights for Mechanistic Interpretability research. The visualizations were generated with the assistance of the OpenAI o4.8 model.

Original post →

More from Research

Research channel →