Switch language한국어
Back to the list

OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding

TL;DR AI

Key summary

2 min read
  1. Researchers introduced OmniPro, the first benchmark for proactive streaming video understanding in omni-modal models.

  2. The benchmark includes 2,700 human-verified samples, 9 sub-tasks, 3 cognitive levels, and modality-isolation labels.

  3. It evaluates models in both probe and online modes to measure when they should respond and what they should say.

  4. Tests on 11 models found audio can help, but is used inconsistently, with performance dropping over time and non-speech audio hardest to handle.

Read the original