Switch language한국어
Back to the list

NVIDIA releases LocateAnything, a high-speed, high-precision object detection AI model that can also detect app UIs and text, not just photos

TL;DR AI

Key summary

2 min read
  1. NVIDIA has released LocateAnything, a fast vision-language model for finding objects in photos, screenshots, and documents.

  2. The model also supports UI element and text detection, and it outperformed Qwen3-VL and Rex-Omni in fine-grained recognition.

  3. A demo is available on Hugging Face, and the model itself has been published openly.

  4. The technology could improve robotics, PC automation, UI analysis, and other real-time perception tasks.

Read the original