Benchmarking Single-Factor Physical Video-to-Audio Generation

TL;DR AI
2 min readKey summary
Researchers introduced FlatSounds, a benchmark for testing whether video-to-audio models capture physical causality from video, not just generate plausible sound.
FlatSounds uses controlled counterfactual pairs and single-video tests to check whether outputs match physical changes and event timing.
Evaluations show current models often rely on captions: captions improve physical and semantic correctness, but can hurt temporal alignment.
The benchmark’s metrics track human judgments well, making it a useful test for trustworthy multimodal generation.
