Paper questions interpretability: Black-box prompting often beats white-box tools
aryaman2020 · x · 2026-08-22
The paper "Pando" investigates if interpretability methods work when models won't explain themselves, introducing a benchmark to control for the "elicitation confounder."
Key Findings:
- Across 720 finetuned models, black-box prompting matches or exceeds white-box methods when explanations are faithful.
- When explanations are absent/misleading, gradient-based attribution improves accuracy by 3-5%, with Relevance Patching (RelP) offering the largest gains.
- Logit Lens, Sparse Autoencoders (SAE), and circuit tracing provided no reliable benefit.
- Gradients track decision computation, while other readouts are dominated by task representation.
More from Research
- FetchMan: Vision-Based Humanoid Policy Trained in Simulation — kevin_zakka · 2026-08-22
- Researcher to publish 20,000-word comprehensive guide to RL for LLMs on Monday — cwolferesearch · 2026-08-22
- Sunday Robotics ACT-2 achieves zero-shot generalization breakthrough — tonyzzhao · 2026-08-22
- WCM: A JEPA-Based World Critic Model Beats SOTA on 149 VLA Robot Tasks — jiqizhixin · 2026-08-22
- Research Note Analyzes Weight Superposition Interference in Neural Networks — thebasepoint · 2026-08-22
- Thought Experiment: Transporting Qwen 27B Weights Back in Time — doodlestein · 2026-08-22