EVL-MCoT: Enhanced Vision-Language Multi-CoT for Harmful Meme Detection

TL;DR AI
2 min readKey summary
Researchers introduced EVL-MCoT, a vision-language framework for harmful meme detection that uses multi-step reasoning across multiple perspectives.
The method combines prototype-guided decoding and context-guided decoding to better align image and text cues, especially for subtle sarcasm and mixed modalities.
Evaluated on HatefulMemes and MultiOff, EVL-MCoT reported strong performance and its code has been made public.
