[Work in Progress]
We release TimeR-v1, a dataset that aims to advance Video LLMs towards temporally grounded reasoning.
We collect existing video-text datasets with timestamp annotations and construct a large-scale high-quality dataset with the help of powerful LLMs.
Each instance in our dataset contains a question-answer pair, as well as a list of timestamps denoting the relevant video segments.