PubMed-OCR is an OCR-centric corpus of scientific articles derived from PubMed Central Open Access PDFs. Each page is rendered to an image and annotated with Google Cloud Vision OCR, released in a compact JSON schema with word-, line-, and paragraph-level bounding boxes.
Scale (release):
This dataset is intended to support layout-aware modeling, coordinate-grounded QA, and… See the full description on the dataset page:
https://huggingface.co/datasets/bevaya/pubmed-ocr.