NVIDIA releases LocateAnything, a high-speed, high-precision object detection AI model that can also detect app UIs and text, not just photos

TL;DR AI
2 min readKey summary
NVIDIA has released LocateAnything, a fast vision-language model for finding objects in photos, screenshots, and documents.
The model also supports UI element and text detection, and it outperformed Qwen3-VL and Rex-Omni in fine-grained recognition.
A demo is available on Hugging Face, and the model itself has been published openly.
The technology could improve robotics, PC automation, UI analysis, and other real-time perception tasks.



