Switch language한국어
Back to the list

DensFiLM: Density-Conditioned Video Saliency for Crowd Scenes

TL;DR AI

Key summary

2 min read
  1. Researchers introduced DensFiLM, a density-conditioned video saliency model for crowd scenes.

  2. It adds a lightweight FiLM module to a Video Swin Transformer to adapt features for sparse versus dense crowds.

  3. On the CrowdFix benchmark, DensFiLM outperformed ACLNet, and the gains came without needing extra optical flow or larger temporal modules.

  4. The results suggest crowd-video saliency improves when the model adjusts to scene density instead of using one fixed attention strategy.

Read the original