Switch language한국어
Back to the list

LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding

TL;DR AI

Key summary

2 min read
  1. Researchers introduced LocateAnything, a unified grounding and detection framework that decodes boxes and points in parallel.

  2. It uses Parallel Box Decoding instead of token-by-token generation, improving inference speed and throughput.

  3. The team also released LocateAnything-Data, a dataset of more than 138 million samples for large-scale training.

  4. Across benchmarks, the system delivers stronger high-IoU localization while keeping fast performance.

Read the original