VideoChat: Based on InternVid, we created additional instruction data and used GPT-4 to condense the existing data.
VideoChatGPT: The original caption data was converted into conversation data based on the same VideoIDs.
Kinetics-710 & SthSthV2: Option candidates were generated from UMTtop-20 predictions.
NExTQA: Typos in the original sentences were corrected.
CLEVRER: For single-option multiple-choice QAs… See the full description on the dataset page:
https://huggingface.co/datasets/wangyueqian/HawkEye-IT.