Switch language한국어
Back to the list

VISAFF: Speaker-Centered Visual Affective Feature Learning for Emotion Recognition in Conversation

TL;DR AI

Key summary

2 min read
  1. Researchers introduced VISAFF, a two-stage, tuning-free framework for speaker-centered emotion recognition in conversation.

  2. It guides frozen vision-language models to focus on the active speaker’s visual affective cues, then adds text and audio when visual evidence is uncertain.

  3. The method uses reliability-guided complementation to reduce overreliance on ambiguous visuals and improve multimodal understanding.

  4. Reported results on two datasets show strong performance, suggesting a more efficient alternative to expensive fine-tuning.

Read the original