RDVSv2: A Large-scale Benchmark for RGB-D Video Salient Object Detection

TL;DR AI
2 min readKey summary
Researchers introduced RDVSv2, a new large-scale RGB-D video salient object detection benchmark built from public stereoscopic online videos.
The dataset includes 249 sequences and 29,077 annotated frames with depth maps and eye-tracking-guided salient object masks.
They also adapted SAM2 with parameter-efficient fine-tuning to jointly use RGB, depth, and optical flow.
The resulting baseline achieved state-of-the-art performance on RDVSv2 and other RGB-D VSOD benchmarks.
RDVSv2 offers a higher-quality testbed for evaluating and improving multi-modal video understanding models.
