When Vision Speaks for Sound
TL;DR AI
2 min readKey summary
Researchers found that many video multimodal models appear to understand audio but actually rely on visual cues, a classic Clever Hans-style failure mode.
The paper introduces Thud, an intervention-based framework using Shift, Mute, and Swap tests to measure audio-visual misalignment.
It also proposes a two-stage training recipe that improves audio verification and raises benchmark performance.
The work matters because it exposes a reliability gap in real-world video AI and offers a practical way to diagnose and reduce it.
