Why SAE Features Fail at Steering: FEGA Framework Explains Downstream Geometry

Tanmoy_Chak · x · 2026-08-01

While Sparse Autoencoder (SAE) features are often interpretable, they frequently fail as stable directions for model steering. A new study introduces FEGA (Feature-Effect Geometry Analysis), a framework to analyze the downstream geometry of these feature effects.

The research categorizes features into two distinct classes:

The analysis reveals that consistent one-dimensional effects are rare across SAE variants, explaining why most features do not behave as reusable steering directions.

Original post →

More from Research

Research channel →