Switch language한국어
Back to the list

Towards Consistent Video Geometry Estimation

TL;DR AI

Key summary

2 min read
  1. Researchers introduced ViGeo, a transformer foundation model for video geometry estimation.

  2. It uses dynamic chunking attention and a completion-based refinement pipeline to learn dense, temporally coherent 3D targets.

  3. ViGeo works in online, offline, and long-video settings, while also predicting surface normals.

  4. The model sets new state-of-the-art results on public benchmarks for depth, normals, and point maps.

Read the original