New benchmark confirms AI video generators look stunning but still can't reason about the world

TL;DR AI
2 min readKey summary
Tsinghua University introduced WorldReasonBench and WorldRewardBench to test whether AI video models can keep scenes physically, socially, logically, and informationally consistent.
The benchmark found commercial models generally outperform open-source systems, with Sora 2, Seedance 2.0, and Veo 3.1-Fast among the stronger performers.
Even the best systems still struggled most with logical reasoning and other grounded world-knowledge tasks.
The results show that realistic-looking video does not necessarily mean a model understands the real world well enough for reliable instruction-following or complex generation.
