VISAFF: Speaker-Centered Visual Affective Feature Learning for Emotion Recognition in Conversation

TL;DR AI
2 min readKey summary
Researchers introduced VISAFF, a two-stage, tuning-free framework for speaker-centered emotion recognition in conversation.
It guides frozen vision-language models to focus on the active speaker’s visual affective cues, then adds text and audio when visual evidence is uncertain.
The method uses reliability-guided complementation to reduce overreliance on ambiguous visuals and improve multimodal understanding.
Reported results on two datasets show strong performance, suggesting a more efficient alternative to expensive fine-tuning.
