Contains 1.5 million document pages from the SafeDocs Common Crawl collection:
https://digitalcorpora.org/corpora/file-corpora/cc-main-2021-31-pdf-untruncated/
Pages are OCRd with word‑level bounding boxes.
Page images have been resized to a maximum dimension of 1024×1024 and are heavily compressed. Bounding-box coordinates are in the original (pre‑resize) image dimensions.
OCR was performed using python-doctr.
Pages have been filtered to keep English and… See the full description on the dataset page:
https://huggingface.co/datasets/albertklorer/safedocs.