A high-quality, human-verified benchmark for evaluating vision-language models on GUI element grounding tasks. Given a screenshot and a natural language query describing an element's functionality, models must localize the target UI element.
FuncElemGnd addresses a critical challenge in GUI understanding: grounding elements by their function rather than appearance. Unlike traditional object detection, this task… See the full description on the dataset page:
https://huggingface.co/datasets/HongxinLi/AutoGUI-v2.0.