AxBench: Simple baselines outperform Sparse Autoencoders in steering LLMs
aryaman2020 · x · 2026-09-02
Stanford et al. released AxBench, the first large-scale benchmark for LLM steering and concept detection. Experiments on Gemma-2-2B and 9B show:
- Steering: Prompting > Finetuning > existing representation methods.
- Concept Detection: Representation methods like DiffMean perform best.
- SAE Performance: Not competitive on both tasks.
The paper also introduces ReFT-r1, a weakly supervised method competitive on both tasks while offering interpretability.
More from Research
- Epoch Index suggests AI capabilities progress twice as fast with reasoning models — Jsevillamol · 2026-09-02
- Study Finds LLM Embeddings Predict Brain Language Response Better Than Simple Metrics — begusgasper · 2026-09-02
- Study Finds AI Interprets Bone Age X-Rays Better Than Clinicians — zakkohane · 2026-09-02
- Study Shows Neural Network Embeddings Approximable by Symbolic Vectors — xuanalogue · 2026-09-02
- otoSpeech Task releases 20 hours of full-duplex task-oriented conversation data — kastnerkyle · 2026-09-02
- Physical AI Deep Dive: From Static Programming to Imitation Learning and Edge Compute — ShawnHymel · 2026-09-02