Removing 10% of 'Important' Cells Cuts Model Output Nearly 3x, New Explainability Test Shows
bravo_abad · x · 2026-09-28
This paper introduces a concrete way to test whether an AI model's highlighted explanations actually drive its predictions: a model predicts clinical outcomes from measured properties of individual tumor cells and their spatial neighbors, an explanation method flags the most important cells, and researchers then remove those cells (with their connections) from the input to compare against removing an equal share of less important cells.
- On a breast-cancer dataset, removing 10% of the important cells caused an average output drop nearly 3x larger than removing the same proportion of unimportant cells
- It also more severely damaged the model's ability to rank patients by survival
- This intervention-based validation offers an empirical way to test explanation methods beyond intuitive visualizations
More from Research
- Neuroscientist pushes back: task-trained nets converge on brain-like representations that can predict primate cortex — coherence · 2026-09-28
- ESPnet's YODAS3 Speech Dataset with 1M+ Entries Trends on Hugging Face — espnet · 2026-09-28
- BoundInk Treats Inter-Character Boundaries as Explicit Units, Cutting DTW by up to 47.8% — SUNGKYUNARCH · 2026-09-28
- TUM's Continuous Depth Batching Unlocks 99% of Speedup for Looped LMs — TUM · 2026-09-28
- ARGUS: An LLM Pipeline Audits Causal Identification Assumptions, Catching 73% of Planted Flaws — Yonghong Zhang · 2026-09-28
- Category theory meets deep learning: PyNCD diagrams derive hardware-aware FlashAttention — GioeleZardini · 2026-09-28