Formal Mechanistic Interpretability: Automated Circuit Discovery with Provable Guarantees

CatAstro_Piyush · x · 2026-08-19

This ICLR 2026 paper introduces formal verification to mechanistic interpretability. The authors propose automated algorithms yielding circuits with provable guarantees, including input domain robustness, robust patching, and minimality. Experiments show significantly stronger robustness guarantees than standard methods.

Original post →

More from Research

Research channel →