We added an "action_description" field to each action, generated by GPT-4, which provides a semantic description of the action.
Semantic Annotation Construction: Following the approach used in the second stage of the AITZ dataset, we designed multimodal inputs for each operation in the Mind2Web dataset. These inputs consist of two screenshots and a prompt:
The first screenshot is the original webpage interface.
The second screenshot shows the original interface with a blue plus sign marking… See the full description on the dataset page:
https://huggingface.co/datasets/xzq11111/Mind2Web_with_act_desc.