This dataset is designed for text and image retrieval tasks. It consists of parsed documents (corpus), generated queries, and relevance judgments (qrels).
The dataset contains three configurations: corpus, queries, and qrels.
Contains the document pages with their text and image content. The images are stored directly within the Parquet files.
corpus_id (string): Unique identifier for the… See the full description on the dataset page:
https://huggingface.co/datasets/eagerworks/multimodal-dataset.