Why SAE Features Fail at Steering: FEGA Framework Explains Downstream Geometry
Tanmoy_Chak · x · 2026-08-01
While Sparse Autoencoder (SAE) features are often interpretable, they frequently fail as stable directions for model steering. A new study introduces FEGA (Feature-Effect Geometry Analysis), a framework to analyze the downstream geometry of these feature effects.
The research categorizes features into two distinct classes:
- Value-like features: Encode concepts themselves (e.g., "Paris") and typically produce structured, low-dimensional effects.
- Pointer-like features: Encode functions operating on context-supplied values (e.g., "copy after match") and predominantly produce diffuse effects.
The analysis reveals that consistent one-dimensional effects are rare across SAE variants, explaining why most features do not behave as reusable steering directions.
More from Research
- Reconstructing >1.5M Cells: Mouse Embryonic Lineage Atlas Released — anshulkundaje · 2026-08-01
- Edison Scientific Launches Hub for Open-Weights Models and Benchmarks — CatAstro_Piyush · 2026-08-01
- Why AI Struggles to Become Einstein: Paper Highlights Lack of Principle Discovery — OdedRechavi · 2026-08-01
- Experiment Shows Training AI Image Models at 512 Resolution + Upscaling Saves VRAM — More_Bid_2197 · 2026-08-01
- MoGe-3 by Microsoft Sets SOTA on 9 Benchmarks for High-Fidelity 3D Geometry from a Single Image — RexDouglass · 2026-08-01
- Experiment Reveals temp=0 Non-determinism Flips LLM Safety Categories — Midoxp · 2026-08-01