CoVR-R is a reasoning-aware benchmark for composed video retrieval. Given a reference video and a textual modification, the goal is to retrieve the correct target video that reflects the requested change and its implied visual consequences.
This dataset is designed for settings where simple keyword overlap is not enough. Many edits require reasoning about state transitions, temporal progression, camera changes, and cause-effect… See the full description on the dataset page:
https://huggingface.co/datasets/omkarthawakar/CoVR-R.