Switch language한국어
Back to the list

RDVSv2: A Large-scale Benchmark for RGB-D Video Salient Object Detection

TL;DR AI

Key summary

2 min read
  1. Researchers introduced RDVSv2, a new large-scale RGB-D video salient object detection benchmark built from public stereoscopic online videos.

  2. The dataset includes 249 sequences and 29,077 annotated frames with depth maps and eye-tracking-guided salient object masks.

  3. They also adapted SAM2 with parameter-efficient fine-tuning to jointly use RGB, depth, and optical flow.

  4. The resulting baseline achieved state-of-the-art performance on RDVSv2 and other RGB-D VSOD benchmarks.

  5. RDVSv2 offers a higher-quality testbed for evaluating and improving multi-modal video understanding models.

Read the original