This comprising 500 complex questions that necessitate cross-modal, multi-hop retrieval to identify the key information. The benchmark contains three types of questions:
Text-Image questions: Questions requiring a relevant image for answering but that image cannot be directly identified by the question
Image-Text questions: Questions requiring a relevant texts for answering but texts cannot be directly identified by the question
Image-Image questions: Simple… See the full description on the dataset page:
https://huggingface.co/datasets/Fujitsu/MM-RAG.