GUIrilla-Gold
Dataset Summary
GUIrilla-Gold benchmark is manually annotated test part from GUIrilla-Task.
Dataset Structure
Data Fields
screen_id
int
Unique screenshot index.
app_name
string
Bundle name (e.g. com.apple.Safari).
task
string
Final, cleaned instruction.
raw_task
string
Raw task draft.
image
image
Full-resolution PNG.… See the full description on the dataset page:
https://huggingface.co/datasets/macpaw-research/GUIrilla-Gold.