Wikimedia Commons Documents
This dataset is created for the evaluation of retrieval models. It contains images of (mostly historic) documents which should be identified based on their description. We extracted those descriptions from Wikimedia Commons. We have included the license type and a link (license_text) to the original Wikimedia Commons page for each extracted image.
The text_description column contains OCR text extracted from the images… See the full description on the dataset page:
https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml_deprecated.