AxBench: Simple baselines outperform Sparse Autoencoders in steering LLMs

aryaman2020 · x · 2026-09-02

Stanford et al. released AxBench, the first large-scale benchmark for LLM steering and concept detection. Experiments on Gemma-2-2B and 9B show:

The paper also introduces ReFT-r1, a weakly supervised method competitive on both tasks while offering interpretability.

Original post →

More from Research

Research channel →