Switch language한국어
Back to the list

Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos

TL;DR AI

Key summary

2 min read
  1. Researchers introduced Artifact-Bench, a benchmark for spotting flaws in AI-generated videos.

  2. It uses a hierarchy of realism artifacts and tests three tasks: real-vs-fake, pairwise realism ranking, and fine-grained artifact ID.

  3. They evaluated 19 leading multimodal LLMs and found widespread weaknesses in judging video realism.

  4. Model judgments often diverged from human preferences, raising concerns for evaluation and safety use cases.

Read the original