This dataset contains a carefully curated ~11B-tokens corpus, which serves as an offline search engine for our data generation process, eliminating the need for external Search APIs. Details on the corpus curation process are available in our blog.
Each row in the dataset contains the⦠See the full description on the dataset page:
https://huggingface.co/datasets/OpenResearcher/OpenResearcher-Corpus.