Switch language한국어
Back to the list

RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video

TL;DR AI

Key summary

2 min read
  1. Researchers introduced RayDer, a unified feed-forward transformer for self-supervised novel view synthesis from unconstrained video.

  2. It combines camera estimation, reconstruction, and rendering in one model, using a minimal dynamic state to cope with changing content during training.

  3. Across model sizes and data scales, RayDer showed strong power-law scaling and competitive zero-shot performance on many benchmarks.

  4. The work points to a simpler, more scalable path for training view-synthesis models on large real-world video datasets without full supervision.

Read the original