OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
TL;DR AI
2 min readKey summary
Researchers introduced OmniPro, the first benchmark for proactive streaming video understanding in omni-modal models.
The benchmark includes 2,700 human-verified samples, 9 sub-tasks, 3 cognitive levels, and modality-isolation labels.
It evaluates models in both probe and online modes to measure when they should respond and what they should say.
Tests on 11 models found audio can help, but is used inconsistently, with performance dropping over time and non-speech audio hardest to handle.
