LVSum: A Benchmark for Timestamp-Aware Long Video Summarization

TL;DR AI
2 min readKey summary
Researchers introduced LVSum, a human-annotated benchmark for timestamp-aware long video summarization.
It covers 72 videos across 13 domains, with up to 10 summaries per video, and evaluates both proprietary and open-source MLLMs.
New LLM-based metrics and standard metrics show that transcripts help more than visual frames, but model summaries still fall well short of human-written ones.
The benchmark reveals persistent weaknesses in temporal grounding, instruction following, and cross-modal coherence.
