Switch language한국어
Back to the list

LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV

TL;DR AI

Key summary

2 min read
  1. Researchers introduced LongAV-Compass, a benchmark for minute-long audio-visual generation across text-, image-, and video-conditioned tasks.

  2. It covers 284 test cases and evaluates T2AV, I2AV, and V2AV with a unified framework using automated and perceptual metrics.

  3. The benchmark measures over 20 dimensions, including quality, consistency, semantic alignment, synchronization, and narrative coherence.

  4. Tests on 11 models include human-alignment checks, helping expose how well systems sustain long-form output quality across input types.

Read the original