OVEarth-Bench: Evaluating Category Breadth and Query Diversity for Open-Vocabulary Earth Observation
TL;DR AI
2 min readKey summary
Researchers released OVEarth-Bench, a new zero-shot benchmark for open-vocabulary earth observation.
It expands evaluation to broader hierarchical categories and multiple natural-language query types.
Tests show overall performance is still limited, with multimodal large language model-based methods doing best.
EO-specific models underperform compared with the top general-purpose approaches, revealing major gaps.
