Human preference dataset for image-to-video (I2V) generation quality. Each row contains a reference image, two generated videos (one from Pika and one from CogVideoX), and 10 human preference annotations aggregated via majority vote.
This is the large (3,000-row) subset — the complete dataset. See also: small (1,000 rows), medium (2,000 rows).