PCA Geometry: Raw Activations vs SAE Feature Strengths
Sauers_ · x · 2026-07-08
Researchers compared the PCA of raw neural network activations (which form a neat circle) with the PCA of Sparse Autoencoder (SAE) feature strengths (which exhibit an irregular shape), revealing significant geometric structural differences. This experiment exposes a potential inconsistency between the SAE feature space and the raw activation space, offering valuable insights for Mechanistic Interpretability research. The visualizations were generated with the assistance of the OpenAI o4.8 model.
More from Research
- Kimi K3 may be strong on cyber, but token efficiency keeps it off UK AISIS — teortaxesTex · 2026-07-27
- ARC AGI 3 should have stayed private, with no examples or public dataset — flowersslop · 2026-07-27
- ExploitGym may have only 60–70% solvable tasks, fueling the OpenAI cheating debate — max_paperclips · 2026-07-27
- Noahpinion quotes Chollet: intelligence may hit a hard ceiling — binarybits · 2026-07-27
- Paper argues graph topology can become the core operating system for AI agents — theomitsa · 2026-07-27
- A question probes how multi-agent branching scales against compute budget and model size — iskander · 2026-07-27