Ground-truth Click & Type actions for macOS screenshots
Dataset Summary
GUIrilla-Task pairs real macOS screenshots with free-form natural-language instructions and precise GUI actions.
Every sample asks an agent either to:
Click a specific on-screen element, or
Type a given text into an input field.
Targets are labelled with bounding-box geometry, enabling exact evaluation of visual-language grounding models.
Data were gathered automatically by… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/GUIrilla-Task.