Rethinking Visual Attribution for Chest X-ray Reasoning in Large Vision Language Models
TL;DR AI
2 min readKey summary
Researchers introduced a causal evaluation framework to test whether chest X-ray vision-language models rely on true image evidence.
They filtered CXR-VQA examples using causal evidence and benchmarked 11 attribution methods across six open-source models.
Many existing attribution methods were found to be unreliable for clinical reasoning and evidence grounding.
To improve this, they proposed MedFocus, a concept-based method that localizes clinically meaningful regions and measures their causal impact.
MedFocus improved explanation quality and trustworthiness for medical vision-language models.
