Towards Consistent Video Geometry Estimation
TL;DR AI
2 min readKey summary
Researchers introduced ViGeo, a transformer foundation model for video geometry estimation.
It uses dynamic chunking attention and a completion-based refinement pipeline to learn dense, temporally coherent 3D targets.
ViGeo works in online, offline, and long-video settings, while also predicting surface normals.
The model sets new state-of-the-art results on public benchmarks for depth, normals, and point maps.
