LocateAnything-Data is the public training-data release for
LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel
Box Decoding.
LocateAnything formulates detection and visual grounding as a unified
vision-language task. Given an image and a category, phrase, text string, or
action-oriented instruction, the model predicts the corresponding bounding
boxes or points. The data spans natural… See the full description on the dataset page:
https://huggingface.co/datasets/NVEagle/LocateAnything-Data.