This repo stores the training datasets used to train the AutoGUI model.
Autogui-625k: This is the v0.1 dataset collected by our AutoGUI annotation pipeline on 05/31/2024. For the v0.2 dataset, refer to
https://huggingface.co/datasets/AutoGUI/AutoGUI-v1-702k
Cauldron: This is one of the two general datasets used to maintain the general visual understanding ability of the trained VLM. We select the Screen2Words, DocVQA, OCR-VQA, visualmrc, infovga, and Diagram image-to-text from the whole… See the full description on the dataset page:
https://huggingface.co/datasets/HongxinLi/AutoGUI-v1-zip.