DensFiLM: Density-Conditioned Video Saliency for Crowd Scenes

TL;DR AI
2 min readKey summary
Researchers introduced DensFiLM, a density-conditioned video saliency model for crowd scenes.
It adds a lightweight FiLM module to a Video Swin Transformer to adapt features for sparse versus dense crowds.
On the CrowdFix benchmark, DensFiLM outperformed ACLNet, and the gains came without needing extra optical flow or larger temporal modules.
The results suggest crowd-video saliency improves when the model adjusts to scene density instead of using one fixed attention strategy.
