ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow

TL;DR AI
2 min readKey summary
ShadowDancer is a new framework for any-action, frame-level control in video world models that learns transferable actions from demonstration videos.
It uses “shadow pairs” — videos with the same dynamics but different appearances — and cross-shadow prediction to learn a unified dynamics representation.
That representation can be reused in new scenes without action labels, motion estimators, or fine-tuning.
The approach could turn demo clips into reusable action assets for more flexible simulation, editing, and interactive generation.
