This dataset is a copy of the original test split from MMTab, taking only items where an 'original_query' is present, and removing the 'input' and 'output' columns, as they are unneccesary for retrieval tasks.
The text_description column contains OCR text extracted from the images using EasyOCR.
This particular dataset is a subsample of 1000 random rows from the full dataset which can be found here.
@misc{zheng2024multimodaltableunderstanding… See the full description on the dataset page:
https://huggingface.co/datasets/jinaai/MMTab.