Switch language한국어
Back to the list

DICA: Dual-Indicator Guided Contrastive Alignment in Multimodal Large Language Models

TL;DR AI

Key summary

2 min read
  1. Researchers proposed DICA, an inference-time method to improve visual grounding in multimodal large language models.

  2. DICA monitors two indicators, Visual Attention Entropy and Output Image Correlation, to detect attention drift or weak image dependence.

  3. When drift is detected, it triggers targeted contrastive alignment to restore grounding.

  4. Across multiple benchmarks, DICA outperformed prior methods and reduced hallucinations.

Read the original