UAI 2026 paper: identifiability metrics show systematic false positives in interpretability evals

RexDouglass · x · 2026-08-19

The paper asks a core question: how can we reliably measure whether a representation identifies latent factors? Can current identifiability metrics actually evaluate interpretability or unsupervised mechanism discovery?

The authors find systematic false positives and false negatives in these evaluations, suggesting many claimed mechanism-discovery results may not be trustworthy.

Original post →

More from Research

Research channel →