Switch language한국어
Back to the list

LVSum: A Benchmark for Timestamp-Aware Long Video Summarization

TL;DR AI

Key summary

2 min read
  1. Researchers introduced LVSum, a human-annotated benchmark for timestamp-aware long video summarization.

  2. It covers 72 videos across 13 domains, with up to 10 summaries per video, and evaluates both proprietary and open-source MLLMs.

  3. New LLM-based metrics and standard metrics show that transcripts help more than visual frames, but model summaries still fall well short of human-written ones.

  4. The benchmark reveals persistent weaknesses in temporal grounding, instruction following, and cross-modal coherence.

Read the original