Self-Supervised Learning of Structured Dynamics from Videos
TL;DR AI
2 min readKey summary
Researchers introduced Structured Dynamics Model, a self-supervised approach that disentangles camera motion from object motion in videos.
The model uses frozen features from a pretrained vision transformer and future-feature prediction to learn structured video dynamics.
Training combines unlabeled video with weak supervision from synthetic Kubric scenes.
On the new ProbeMotion benchmark, it outperformed simple backbone features and competed well with some strongly supervised methods.
