LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
TL;DR AI
2 min readKey summary
Researchers introduced LocateAnything, a unified grounding and detection framework that decodes boxes and points in parallel.
It uses Parallel Box Decoding instead of token-by-token generation, improving inference speed and throughput.
The team also released LocateAnything-Data, a dataset of more than 138 million samples for large-scale training.
Across benchmarks, the system delivers stronger high-IoU localization while keeping fast performance.
