PCA Geometry: Raw Activations vs SAE Feature Strengths
Sauers_ · x · 2026-07-08
Researchers compared the PCA of raw neural network activations (which form a neat circle) with the PCA of Sparse Autoencoder (SAE) feature strengths (which exhibit an irregular shape), revealing significant geometric structural differences. This experiment exposes a potential inconsistency between the SAE feature space and the raw activation space, offering valuable insights for Mechanistic Interpretability research. The visualizations were generated with the assistance of the OpenAI o4.8 model.
More from Research
- LoMa Paper Ships REALLY HardPairs Dataset, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Johns Hopkins Launches Full-Stack Hands-on Robot Learning Class with SO-101 Arm Kits — _krishna_murthy · 2026-09-11
- SyncWorld: In-Context Robot World Model Simulates Unseen Views and Embodiments Zero-Shot — ChongZzZhang · 2026-09-11
- A 3D Pose Dataset for Dogs Released — ducha_aiki · 2026-09-11
- Five tells that still make AI video read as AI, from physics glitches to missing operators — NewPhoneWhotiz · 2026-09-11
- AnyMatch accepted to ECCV 2026 with a NoPresenter design — ducha_aiki · 2026-09-11