TimeWarp is a multimodal synthetic temporal preference data generation pipeline for enhancing
temporal understanding in Video Large Language Models (Video-LLMs).
It focuses on understanding event order, temporal commonsense, and
implicit sequence relationships in multimodal (video + text) contexts.
Modality: Video + Text
Goal: Measure and improve a modelβs ability to understand temporal dynamics in visual scenes
Format: Video frames /β¦ See the full description on the dataset page:
https://huggingface.co/datasets/time-warp/timewarp.