We compile a large set of evaluation tasks to understand the capabilities of multimodal embedding models. This benchmark covers 4 meta tasks and 36 datasets meticulously selected for evaluation.
The dataset is published in our paper VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks.
For each dataset, we have 1000 examples for evaluation. Each example contains a query and a set of… See the full description on the dataset page:
https://huggingface.co/datasets/TIGER-Lab/MMEB-eval.