Video-language pair classification dataset from VideoCon (Bansal et al., 2023).
Each row contains a video and a text caption. Label=1 means the caption correctly describes the video;
label=0 means it is a semantically-plausible contrast caption (entity/action/attribute swaps, event order flips).
Source: videocon/videocon (videocon_human.csv) — 569 videos from ActivityNet, 1138 pairs total (2 corrupt videos removed).