This dataset is derived from the Slideshare-1M dataset, using only the subset under 8 of the full dataset for the test set creation.
Queries are manually annotated by two annotators.
The text_description column contains OCR text extracted from the images using EasyOCR.
This particular dataset is a subsample of 1000 rows from the full dataset which can be found here.
This dataset may contain publicly available images or text data. All… See the full description on the dataset page:
https://huggingface.co/datasets/jinaai/stanford_slide.