LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV
TL;DR AI
2 min readKey summary
Researchers introduced LongAV-Compass, a benchmark for minute-long audio-visual generation across text-, image-, and video-conditioned tasks.
It covers 284 test cases and evaluates T2AV, I2AV, and V2AV with a unified framework using automated and perceptual metrics.
The benchmark measures over 20 dimensions, including quality, consistency, semantic alignment, synchronization, and narrative coherence.
Tests on 11 models include human-alignment checks, helping expose how well systems sustain long-form output quality across input types.
