MEA: multi-agent reward optimization lifts ML explanation faithfulness by up to 34%

UVABiology · hf · 2026-10-06

Researchers propose MEA, a multi-agent framework that removes the expertise barrier to ML explainability. A Proposer agent selects and configures explanation tools based on question and modality, while an Actor agent is optimized end-to-end against faithfulness to produce natural-language explanations grounded in model behavior across tabular, text, and vision data. The work introduces question types spanning feature attribution, counterfactual reasoning, and spurious feature detection, each with perturbation-based faithfulness metrics. Notably, frontier LLMs are found to systematically produce unfaithful explanations. With reward-driven optimization plus a modality-adaptive penalty, MEA beats post-hoc, agentic, and closed-source baselines across six datasets, gaining +28% (tabular), +21% (text), and +34% (vision) faithfulness over the untrained backbone.

Original post →

More from coding & agent

coding & agent channel →