2 examples from 1 videos - UI automation tasks from screen recordings.
Dataset Structure
Each entry contains:
video_id: Sequential ID for each video (video_001, video_002, etc.)
step: Step number within that video (0, 1, 2, ...)
system: System prompt for the GUI agent
user: Task instruction + previous actions
assistant: Model's reasoning and action
image: Screenshot of the UI state