The training data split used for training VLM2Vec models in the paper VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks (ICLR 2025).
MMEB benchmark covers 4 meta tasks and 36 datasets meticulously selected for evaluating capabilities of multimodal embedding models.
During training, we utilize 20 out of the 36 datasets.
For evaluation, we assess performance on the 20 in-domain (IND) datasets and the remaining 16… See the full description on the dataset page:
https://huggingface.co/datasets/Y-J-Ju/MMEB-train.