HARD-TIME evaluates whether video-language models can localize moments in time and avoid answers that are
not supported by the video. This repository contains benchmark annotations for the evaluation tasks; it does
not include source videos, transcripts, or evidence artifacts. Access to the annotations does not grant any
license or reuse rights for the underlying third-party content.