This dataset is derived from the CharXiv dataset, reformatting the test split with modified field names, so that it can be used in the ViDoRe benchmark.
The text_description column contains OCR text extracted from the images using EasyOCR.
This particular dataset is a subsample of 1000 random rows from the full dataset which can be found here.
This dataset may contain publicly available images or text data. All data is provided for research… See the full description on the dataset page:
https://huggingface.co/datasets/jinaai/CharXiv-en_deprecated.