This dataset is a version of the peS2o dataset restricted to openly licensed articles.
Pes2o is derived from S2ORC, a corpus of openly licensed abstract and full-text papers that have been converted to a structured format using Grobid.
Starting from Grobid’s XML output, peS2o filters papers that are too short, have incorrect metadata, are in languages other than English, and contain OCR errors using a combination of heuristic- and model-based filtering… See the full description on the dataset page:
https://huggingface.co/datasets/common-pile/peS2o.