Mechanistic Interpretability Explained: Linear Probes, Feature Maps, and Why AIs Evade Them

Astral Codex Ten · rss · 2026-09-08

Scott Alexander publishes a long-form guide to mechanistic interpretability—the science of "reading an AI's mind"—covering the field's history, key techniques, and core frustrations.

The story so far

Linear probes

Why concept suppression backfires

Positioned as the author's personal reference log, with more techniques to come in later installments.

Original post →

More from Safety

Safety channel →