Interpretability Tools Fail to Beat Reading Transcripts in Agent Diagnostics

a_karvonen · x · 2026-08-22

A study tested three activation-based tools—activation oracles, natural language autoencoders, and SAEs—that had succeeded in prior auditing games. None of them outperformed simply reading the transcript in diagnosing agent behavior. This result remained robust across hyperparameters, prompting strategies, and Fable goal loops.

Related event: Interpretability tools fall short of just reading the transcript(2 posts)→

Original post →

More from Safety

Safety channel →