This dataset is a subset consisting of 1000 records from the originally-sized 360GB wiki-ss-corpus (Wiki Screenshot corpus) dataset that was created by the DSE model authors. The dataset consists of scraped screenshots of Wikipedia pages, some curation was done to the screenshots (to ensure dataset quality), and then a few metadata points such as the document id were given to each record.
features:
- name: image
dtype: image
- name:… See the full description on the dataset page: https://huggingface.co/datasets/andreaparker/wiki-ss-corpus-train-sm-subset.