Do Audio-Visual Large Language Models Really See and Hear?
TL;DR AI
2 min readKey summary
Researchers performed the first mechanistic interpretability study of Audio-Visual Large Language Models (AVLLMs).
They found rich audio semantics encoded in intermediate layers, but these signals often fail to influence final text when audio conflicts with vision.
Probing showed deeper fusion layers disproportionately favor visual representations, and the imbalance traces to limited audio-specific alignment during training.
