Paper questions interpretability: Black-box prompting often beats white-box tools

aryaman2020 · x · 2026-08-22

The paper "Pando" investigates if interpretability methods work when models won't explain themselves, introducing a benchmark to control for the "elicitation confounder."

Key Findings:

Related event: Studies Question Interpretability Tools: NLA Found Impractical and Misleading(3 posts)→

Original post →

More from Research

Research channel →