Computer use trajectories for training Vision-Language-Action models.
image: Screenshot (PIL Image)
instruction: Task description
action_type: 0=MOUSE, 1=KEYBOARD
mouse_x, mouse_y: Normalized coordinates [0,1]
click_type: 0-8 (NO_CLICK, LEFT_CLICK, etc.)
keyboard_text: Text with special tokens
os_type: ubuntu, windows_macos
episode_id, step_idx: Episode structure
Index
Type
Description
0
NO_CLICK… See the full description on the dataset page:
https://huggingface.co/datasets/TESS-Computer/tess-agentnet.