Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos
TL;DR AI
2 min readKey summary
Researchers introduced Artifact-Bench, a benchmark for spotting flaws in AI-generated videos.
It uses a hierarchy of realism artifacts and tests three tasks: real-vs-fake, pairwise realism ranking, and fine-grained artifact ID.
They evaluated 19 leading multimodal LLMs and found widespread weaknesses in judging video realism.
Model judgments often diverged from human preferences, raising concerns for evaluation and safety use cases.
