The dataset used to train and evaluate ReT for multimodal information retrieval. The dataset is almost the same as the original M2KR, with a few modifications:
we exlude any data from MSMARCO, as it does not contain query images;
we add passage images to OVEN, InfoSeek, E-VQA, and OKVQA. Refer to the paper for more details.
Sources
! Update 12/09/2025
We have just released ReT-2: Recurrence Meets Transformers for Universal Multimodal Retrieval
Download images… See the full description on the dataset page: https://huggingface.co/datasets/aimagelab/ReT-M2KR.