The peS2o dataset is a collection of ~40M creative open-access academic papers,
cleaned, filtered, and formatted for pre-training of language models. It is derived from
the Semantic Scholar Open Research Corpus(Lo et al, 2020), or S2ORC.
We release multiple version of peS2o, each with different processing and knowledge cutoff
date. We recommend you to use the latest version available.
If you use this dataset, please cite:
@techreport{peS2o,
author =… See the full description on the dataset page:
https://huggingface.co/datasets/allenai/peS2o.