Quantitative Video World Model Evaluation for Geometric-Consistency

TL;DR AI
2 min readKey summary
Researchers propose PDI-Bench and PDI-Dataset to quantitatively evaluate whether generated videos preserve 3D geometry.
The framework combines segmentation, point tracking, and monocular reconstruction to test scale-depth alignment, 3D motion consistency, and structural rigidity.
It is designed to expose geometry failures in video world models that subjective or perceptual metrics can miss.
The benchmark uses tools such as SAM 2, MegaSaM, and CoTracker3 to assess perspective distortion in generated clips.
