Vision-Language Models (VLMs) have enabled computer use agents (CUAs) that operate GUIs autonomously with great potential.
However, developing robust CUAs requires extensive in-domain knowledge about software interfaces and operations.
Unlike image–text pairs that are widely available on the Internet, computer-use data, particularly… See the full description on the dataset page:
https://huggingface.co/datasets/OpenGVLab/ScaleCUA-Data.