Rethinking Visual Attribution for Chest X-ray Reasoning in Large Vision Language Models

TL;DR AI
2 min readKey summary
Researchers evaluated causal visual attribution for chest X-ray VQA across 11 methods, six open-source large vision-language models, and two answer modes.
They found that existing attribution methods often fail to identify the evidence the models actually use, exposing a trust gap in medical AI explanations.
The team introduced MedFocus, which highlights clinically meaningful regions and estimates their causal effect more reliably.
The work aims to verify whether LVLM answers are truly grounded in the chest X-ray evidence clinicians care about.
