ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow
TL;DR AI
2 min readKey summary
ShadowDancer is a new framework that turns demonstration videos into reusable action representations for frame-level control of video world models.
It learns transferable dynamics by pairing videos with the same motion but different appearances, then predicting one view from the other through cross-shadow prediction.
This lets ordinary videos serve as action supervision without action labels, motion estimation, or fine-tuning.
The method outperforms strong baselines on action transfer and longer rollouts.
Overall, it could broaden how interactive video models are trained and controlled across diverse environments.
