Switch language한국어
Back to the list

Internalizing Temporal Consistency in Video Object-Centric Learning without Explicit Regularization

TL;DR AI

Key summary

2 min read
  1. Researchers introduced xSSC for video object-centric learning, replacing slot-slot contrastive learning with built-in temporal modeling.

  2. Chrono-Channel Decomposition separates static and dynamic slot information.

  3. Cross-Temporal Reconstruction trains the model to reconstruct features across adjacent frames using a standard reconstruction loss.

  4. The method removes an extra contrastive objective, improving efficiency and achieving new results on object discovery and recognition benchmarks.

Read the original