Switch language한국어
Back to the list

RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video

TL;DR AI

Key summary

2 min read
  1. RayDer is a unified feed-forward transformer for self-supervised novel view synthesis.

  2. It combines camera estimation, scene reconstruction, and rendering in one model, with a minimal dynamic state for changing content.

  3. The system trains stably on unconstrained real-world videos and scales predictably with more data and larger models.

  4. It delivers competitive zero-shot results across many benchmarks, pointing to better 3D understanding without costly labels.

Read the original